Beacon hunter
Malware that phones home leaves a rhythm in the traffic: a check-in every so often, hour after hour. Defenders call it beaconing, and the whole game is telling that rhythm apart from the many benign things that are also periodic. Six exhibits take it from the raw flow log to a tuned rule and a write-up, using the maths of timing, jitter and frequency, on one synthetic hour of traffic with three beacons hidden in it.
Everything runs in this tab and nothing is uploaded. The hour is invented and rebuilt from a fixed seed, so every visitor sees the same traffic: about 40 internal hosts on 10.0.0.0/8, roughly 6,600 flow records and 4,000 DNS queries. Every outside address is from a documentation block (192.0.2, 198.51.100, 203.0.113) and every name ends in .example, .test or .invalid. No malware, no implant, no destination you could reach: only records of connections that never happened.
Lab exploredYou have read the haystack, measured a pair's rhythm, taken it into the frequency domain, weighed size and DNS alongside timing, tuned a transparent rule against the innocent periodic traffic it must leave alone, and written the finding up with a technique and a response. The Entropy lab is a sibling: the same idea of a measurable signal, applied to how random a file looks.
One hour of flows
A flow record is one connection reduced to a line: when it started, who talked to whom, on what port, and how many bytes each way. No payload, no content. An hour of a small network is thousands of these. Raw volume is a poor way in, because a handful of downloads dwarf everything else, so sort by counts and rank instead.
The timeline below plots one dot per connection: time runs left to right across the hour, and each row is one destination, the busiest at the bottom. A periodic source draws a row of evenly spaced dots, a railroad track. Human browsing scatters. Pick a destination to light up its track and read its flows.
| # | Destination | Conns | Bytes | Hosts |
|---|
The maths
A flow is a tuple: (t, source, destination, port, protocol, bytes up, bytes down, duration). It records that a connection happened and how big it was, never what was said. Two assumptions ride underneath: the clock is roughly right (times are seconds from the start of the window, and a beacon's period is only as sharp as the clock), and each connection is one record (a long-lived flow that the collector cut into several will read as several).
Why raw volume misleads. Transfer sizes are heavy-tailed: a few large downloads carry most of the bytes, so a sort by bytes buries a small, chatty beacon that moves almost nothing. Counts and rank are steadier. The busiest destination by connections is rarely the busiest by bytes.
rank by count, not sum · a beacon is small and frequent, a download is large and rare
In practiceThis is Zeek's conn.log, or NetFlow/IPFIX from a router, or a proxy log. The fields differ but the shape is the same, and the first move is always the same: aggregate by (source, destination, port) and rank, rather than reading lines.
DefenceKeep flow logs even when you cannot keep payload. Beaconing is a property of when connections happen, which flow records preserve perfectly and cheaply, for a long retention window.
The gaps between connections
Pick a pair and look at the gaps between its connections. Browsing gives a ragged spread of gaps; a beacon gives nearly the same gap over and over. One number captures it: the coefficient of variation, CV = σ / μ, the standard deviation of the gaps over their mean. A memoryless process sits at CV near 1; a metronome sits near 0. A beacon that adds jitter sits in between, and that is where the interesting cases live.
The histogram and strip below are this pair's gaps. The panel reads the statistics the engine computes live: mean, standard deviation, CV, median and MAD, the largest gap and the jitter.
Ranked by connection count. The three beacons are in here somewhere.
- Connections
- 0
- Mean gap
- 0
- Std dev
- 0
- CV = σ/μ
- 0.000
- Median / MAD
- 0
- Largest gap
- 0
The maths
Gaps from a memoryless (Poisson) source are exponentially distributed. An exponential has its standard deviation equal to its mean, so CV = 1. A fixed-period beacon has the same gap every time, so σ = 0 and CV = 0.
Jitter. Give a beacon a period P and multiply each interval by 1 + u, with u uniform on [−j, j]. A uniform variable on a width of 2jP has variance (2jP)2/12 = j2P2/3, so σ = jP/√3 and the mean is P:
CV = j / √3
So a beacon with ±30% jitter reads CV ≈ 0.30/1.732 ≈ 0.17. A detector that flags “CV below 0.1” catches the metronome and misses this one. The unit tests generate a large jittered sample and check the engine’s CV against j/√3.
Median and MAD resist outliers. One missed check-in doubles a gap and inflates σ, so the mean-based CV over-reacts. The median gap and the median absolute deviation from it (MAD) ignore a single stray gap. And the CV estimate itself has error: for near-normal gaps its standard error is about CV/√(2(n−1)), so a handful of samples cannot pin a CV down. The tests check that formula by resampling.
In practiceThis is the core of most flow-based beacon hunts: group by (source, destination), take the gaps, and score the regularity. RITA and many SIEM rules do a version of this, often with the median and MAD rather than the mean, for exactly the outlier reason above.
DefenceScore regularity, not just volume, and use a robust statistic. But do not stop at CV: benign automation is regular too, which the next exhibits make painfully clear.
Into the frequency domain
Bin a pair's connections into a count per second and you have a signal. Two tools read a period out of it. The autocorrelation asks how much the signal looks like itself shifted by each lag: a periodic source shows a comb of peaks at the period and its multiples. The amplitude spectrum (a fast Fourier transform, implemented here) shows a spectral line at one over the period, with harmonics. Jitter smears that line into a hump; a Poisson source has no line at all, just a flat spectrum.
Pick a pair to read its period. The slider adds jitter to a clean copy of the fixed-interval beacon so you can watch its spectral line decay and vanish.
- Dominant period
- n/a
- Peak-to-median
- 0
- Samples
- 0
The maths
The discrete Fourier transform of a length-N signal is
X[k] = ∑n=0N−1 x[n] e−2πi kn/N
The fast Fourier transform computes it in O(N log N). A window of T seconds resolves frequencies to 1/T, and the highest frequency it can see is the Nyquist limit, one every two bins. A single spike splits its energy across bins unless it sits on one exactly, which is spectral leakage; a Hann window tapers the ends and narrows that spread at the cost of a slightly wider main lobe.
Autocorrelation and the spectrum are the same thing. The Wiener-Khinchin relation says the autocorrelation is the inverse transform of the power spectrum |X[k]|2. The page computes the autocorrelation that way and the tests check it equals the direct sum to a part in 109. Parseval’s identity, ∑|x[n]|2 = (1/N)∑|X[k]|2, is checked too.
Why jitter kills the line. Jitter of ±a around a grid multiplies the spectral line by the characteristic function of the jitter, a sinc. With a = fP as a fraction f of the period, the line at the fundamental falls by |sin(2πf) / (2πf)|, reaching zero when the jitter spans a full period. So a jittered beacon has a low peak-to-median ratio and hides from a pure spectral detector, even though its CV still gives it away. The tests compare the measured line height with this sinc model.
Too few samples. Three connections in an hour is two gaps. No spectral method can call that periodic: the resolution is 1/T and there is barely a cycle to see. The page refuses to claim a period below a minimum sample count, and says so.
In practiceFrequency analysis is a strong second opinion when timing looks regular but you want the period, or when you suspect a beacon under noise. It is weak on jittered or low-rate beacons, which is why it sits beside the CV, not instead of it.
DefenceUse the spectrum to confirm and measure a period, not to find every beacon. Expect an attacker who knows this to jitter, which trades a clean spectral line for a slightly higher CV: the two views cover each other.
Timing is not the only tell
A beacon usually sends the same small request each time, so its bytes-up are as regular as its timing, and often the reply is small too. That size regularity survives jitter: an attacker can randomise when without changing what. And DNS carries its own tell. A tunnel sends many queries with long, high-entropy labels under one parent domain, often as TXT records, because that is how you push bytes out over a channel that is hard to block.
Two tables below: pairs by their size regularity, and DNS by source and parent domain. Find the beacon whose timing is ragged but whose size never moves, and find the host that is talking in noise.
| Pair | Conns | Timing CV | Size CV | Up:Down | Same size |
|---|
| Source | Parent | Queries | Unique labels | Bits / label | TXT |
|---|
The maths
Size regularity is the CV again, taken over the bytes-up of each connection rather than the gaps. A beacon's request is templated, so its size CV is tiny even when its timing CV is not. The fraction of connections with an identical byte count is a blunter version of the same fact.
DNS label entropy. A label of L characters drawn uniformly from an alphabet of A symbols carries
H = L log2 A bits
base32 labels give 5 bits per character, hex 4, English words only about 2 to 3. So a 40-character base32 label carries about 200 bits, more than any real hostname needs. The page measures the Shannon entropy of each label's own characters and multiplies by its length; the tests check that a uniform base32 or hex string lands near L log2 A.
The channel ceiling. A DNS label is at most 63 bytes and a full name at most 253, so each query smuggles a few dozen bytes at most. That makes DNS a slow channel, but a resilient one: it rides the resolver every network already trusts, which is why the tell is the shape of the traffic (entropy, length, count, TXT share), not any one query.
In practiceSize and DNS features are what let a hunt separate a beacon from benign automation and catch a channel that has no fixed period at all. Detectors such as Zeek's DNS scripts flag long labels, high entropy and TXT-heavy traffic under one domain.
DefenceWatch bytes-up regularity and DNS label entropy alongside timing. A well-known, popular destination reached by many hosts is almost never a beacon, however regular any one host's connection to it looks: popularity is a strong prior.
Build a rule you can defend
Combine the features into one transparent score per pair, weighted the way you choose, and rank. The truth is hidden until you press Score my rule; then the panel shows precision, recall and F1, and a precision-recall curve. The hard part is not catching the beacons, it is the benign traffic that is just as periodic: the monitoring agent, NTP, the printer, the mail client. Clear each with the allowlist, and read what a real attacker hiding behind that class would get away with.
| # | Pair | Score | CV | PMR | n | Verdict |
|---|
The maths
Score every pair by a weighted sum of features, each scaled so that more beacon-like reads higher. Because the score is linear in the weights, raising a weight never lowers a pair that has more of that feature than another: the tests check that monotonicity. A minimum-count gate drops pairs too short to judge.
precision = TP / (TP + FP), recall = TP / (TP + FN), F1 = 2·P·R / (P + R)
Why precision is hard here. Beacons are rare. With three real beacons among about 1,500 pairs, a false-positive rate of just 1% flags roughly one per cent of the innocent pairs, which is far more than three:
false alarms per beacon ≈ FPR · (pairs) / (beacons)
so even a good-looking rule drowns the true positives unless you cut the candidate pool. That is what the allowlist does: it removes a class of benign traffic (and states the cost of doing so) rather than lowering the bar. The precision-recall curve and its average precision are computed exactly from the ranked list, and the tests check them against a brute-force count.
In practiceReal beacon detection tunes on a corpus that already contains the network's own automation, counts false positives first, and treats a high score as a lead, not a verdict. The allowlist is a real control, with a real cost: every class you wave through is a class an attacker can hide in.
DefencePrefer few, explainable features over a black box you cannot tune. Publish the cost of each allowlist entry. Revisit the rule when the network's automation changes, because your false-positive floor moves with it.
Turn it into something a team can act on
A finding is only useful written down. Pick the suspected beacon; the page fills what the engine measured. Choose the technique it maps to, pick a response plan that contains no wrong move, and the page generates a plain-text description and a rule sketch you can copy. It also spells out what would fool the rule, so you see the evasion surface before an attacker does.
What would fool this rule
The maths
No new maths here: the report reads the numbers exhibits II to V measured (period, jitter, count, size, first and last seen) and lays them out. The tests check that the fields in the generated text equal the engine’s numbers, and that the technique the page accepts for each beacon is the right one.
Where this sits. Flow-based beaconing detection runs on Zeek conn.log, NetFlow/IPFIX or proxy logs, upstream of full packet capture. It is cheap, it keeps for a long window, and it survives encryption, because it never needs the payload. The rule sketch here is a teaching artefact: a real deployment needs tuning, an allowlist and a review, exactly as exhibit V argued.
In practiceA good beacon write-up names the pair, the measured period and jitter, the technique, and a response with an order to it: contain, then hunt for the same interval and destination elsewhere, then remediate. It hands the next analyst a repeatable query, not a screenshot.
DefenceDetect, then hunt sideways: the same interval and destination on other hosts is the difference between one infection and an incident. Rebooting the host first destroys the evidence and tips off the operator; contain at the network instead.