Everything runs in this tab and nothing is uploaded. The page rebuilds the hour from a fixed seed (about 40 hosts, 6,600 flows and 4,000 DNS queries), every outside address comes from a documentation block, and every name ends in .example, .test or .invalid.
Beacon hunter
Malware that phones home leaves a rhythm in the traffic, a check-in every so often, and the hard part is telling it apart from benign traffic that's periodic too. Six exhibits take one synthetic hour with three hidden beacons from the raw flow log to a tuned rule and a write-up.
Lab exploredThat's the whole hunt, from raw log to write-up. The Entropy lab measures a different signal the same way: how random a file looks.
One hour of flows
A flow record is one connection reduced to a line: when it started, who talked to whom, on what port, and how many bytes each way. No payload, no content. An hour of a small network is thousands of these. Raw volume is a poor way in, because a handful of downloads dwarf everything else. Count connections and rank instead.
The timeline below plots one dot per connection: time runs left to right across the hour, and each row is one destination, the busiest at the bottom. A periodic source draws a row of evenly spaced dots, a railroad track. Human browsing scatters. Pick a destination to light up its track and read its flows.
| # | Destination | Conns | Bytes | Hosts |
|---|
The maths
A flow is a tuple: (t, source, destination, port, protocol, bytes up, bytes down, duration). It records that a connection happened and how big it was, never what was said. It assumes the clock is roughly right and that each connection is one record, so a long flow the collector split will read as several.
Why raw volume misleads. Transfer sizes are heavy-tailed: a few large downloads carry most of the bytes, so a sort by bytes buries a small, chatty beacon that moves almost nothing. Counts and rank are steadier. The busiest destination by connections is rarely the busiest by bytes.
rank by count, not sum · a beacon is small and frequent, a download is large and rare
In practiceThis is Zeek's conn.log, or NetFlow/IPFIX from a router, or a proxy log. The fields differ but the shape is the same, and so is the first move: aggregate by (source, destination, port) and rank, instead of reading lines.
DefenceKeep flow logs even when you can't keep payload. Beaconing is a property of when connections happen, and flow records preserve that perfectly and cheaply for a long retention window.
The gaps between connections
Pick a pair and look at the gaps between its connections. Browsing is all over the place. A beacon checks in after roughly the same gap every time. The number for that is the coefficient of variation, CV = σ / μ: about 1 for random, memoryless arrivals, about 0 for a metronome, and somewhere in between once the operator adds jitter.
The histogram and strip below show this pair's gaps. The panel reads out the stats the engine computes live: mean, standard deviation, CV, median and MAD, the largest gap and the jitter.
Ranked by connection count. The three beacons are in here somewhere.
- Connections
- 0
- Mean gap
- 0
- Std dev
- 0
- CV = σ/μ
- 0.000
- Median / MAD
- 0
- Largest gap
- 0
The maths
Gaps from a memoryless (Poisson) source are exponentially distributed. An exponential's standard deviation equals its mean, so CV = 1. A fixed-period beacon has the same gap every time, so σ = 0 and CV = 0.
Jitter. Give a beacon a period P and multiply each interval by 1 + u, with u uniform on [−j, j]. A uniform variable on a width of 2jP has variance (2jP)2/12 = j2P2/3, so σ = jP/√3 and the mean is P:
CV = j / √3
So a beacon with ±30% jitter reads CV ≈ 0.30/1.732 ≈ 0.17. A detector that flags “CV below 0.1” catches the metronome and misses this one. The unit tests generate a large jittered sample and check the engine’s CV against j/√3.
Median and MAD resist outliers. One missed check-in doubles a gap and inflates σ, while the median gap and the median absolute deviation (MAD) ignore it. The CV estimate has error too: for near-normal gaps its standard error is about CV/√(2(n−1)), so a handful of samples can't pin it down.
In practiceThis is the core of most flow-based beacon hunts: group by (source, destination), take the gaps, score the regularity. RITA and many SIEM rules do a version of this, often with the median and MAD instead of the mean, for the outlier reason above.
DefenceScore regularity as well as volume, with a statistic that shrugs off outliers (median and MAD). Benign automation is regular too, so CV alone never settles it.
Into the frequency domain
Bin a pair's connections into a count per second and you have a signal. Its autocorrelation shows a comb of peaks at the period and its multiples, and its amplitude spectrum (an FFT written for this page) shows a line at one over the period. Jitter smears that line into a hump, and a Poisson source has no line at all.
Pick a pair to read its period. The slider adds jitter to a clean copy of the fixed-interval beacon so you can watch its spectral line decay and vanish.
- Dominant period
- n/a
- Peak-to-median
- 0
- Samples
- 0
The maths
The discrete Fourier transform of a length-N signal is
X[k] = ∑n=0N−1 x[n] e−2πi kn/N
The fast Fourier transform computes it in O(N log N). A window of T seconds resolves frequencies to 1/T, up to the Nyquist limit of one every two bins. A line that falls between bins spreads its energy (spectral leakage), and a Hann window narrows that spread at the cost of a slightly wider main lobe.
Autocorrelation and the spectrum are the same thing. The Wiener-Khinchin relation says the autocorrelation is the inverse transform of the power spectrum |X[k]|2. The page computes the autocorrelation that way, and the tests check it equals the direct sum to a part in 109. The tests check Parseval’s identity, ∑|x[n]|2 = (1/N)∑|X[k]|2, as well.
Why jitter kills the line. Jitter of ±a around a grid multiplies the spectral line by a sinc. With a = fP, the line at the fundamental falls by |sin(2πf) / (2πf)| and reaches zero when the jitter spans a full period. A jittered beacon can slip past a spectral detector while its CV still flags it.
Too few samples. Three connections in an hour is two gaps. No spectral method can call that periodic: the resolution is 1/T and there is barely a cycle to see. The page refuses to claim a period below a minimum sample count, and says so.
In practiceFrequency analysis is a strong second opinion when timing looks regular and you want the period. It's weak on jittered or low-rate beacons, so run it beside the CV.
DefenceUse the spectrum to confirm and measure a period. An attacker who jitters trades a clean spectral line for a higher CV, so the two views cover each other.
Other ways to spot a beacon
A beacon usually sends the same small request each time, so its bytes-up are as regular as its timing. That survives jitter, because randomising when doesn't change what. DNS has a pattern of its own: a tunnel sends many long, high-entropy labels under one parent domain, often as TXT records.
Two tables below: pairs by their size regularity, and DNS by source and parent domain. Find the beacon whose timing is ragged but whose size never moves, and find the host that is talking in noise.
| Pair | Conns | Timing CV | Size CV | Up:Down | Same size |
|---|
| Source | Parent | Queries | Unique labels | Bits / label | TXT |
|---|
The maths
Size regularity is the CV again, taken over the bytes-up of each connection rather than the gaps. A beacon's request is templated, so its size CV is tiny even when its timing CV is not. The fraction of connections with an identical byte count is a blunter version of the same fact.
DNS label entropy. A label of L characters drawn uniformly from an alphabet of A symbols carries
H = L log2 A bits
base32 labels give 5 bits per character, hex 4, and English words only about 2 to 3. A 40-character base32 label carries about 200 bits, more than any real hostname needs. The page multiplies each label's Shannon entropy per character by its length.
The channel ceiling. A DNS label is at most 63 bytes and a full name at most 253, so each query carries a few dozen bytes. DNS is slow, but it rides the resolver every network trusts, so you spot it by the shape of the traffic: entropy, length, count and TXT share.
In practiceSize and DNS features are what let a hunt separate a beacon from benign automation and catch a channel that has no fixed period at all. Detectors such as Zeek's DNS scripts flag long labels, high entropy and TXT-heavy traffic under one domain.
DefenceWatch bytes-up regularity and DNS label entropy alongside timing. A well-known, popular destination reached by many hosts is almost never a beacon, however regular any one host's connection to it looks: popularity is a strong prior.
Build a rule you can defend
Combine the features into one score per pair, weight them your way, and rank. Score my rule reveals precision, recall, F1 and a precision-recall curve. The hard part is the benign traffic that's just as periodic (monitoring, NTP, the printer, mail). Clear each with the allowlist and see what an attacker hiding behind it would get away with. the printer is easily the most punctual thing on the network lol.
| # | Pair | Score | CV | PMR | n | Verdict |
|---|
The maths
Score every pair by a weighted sum of features, each scaled so that more beacon-like reads higher. The score is linear in the weights, so raising a weight never lowers a pair that has more of that feature than another, and the tests check that monotonicity. A minimum-count gate drops pairs too short to judge.
precision = TP / (TP + FP), recall = TP / (TP + FN), F1 = 2·P·R / (P + R)
Why precision is hard here. Beacons are rare. With three among about 1,500 pairs, a 1% false-positive rate flags far more innocent pairs than real ones:
false alarms per beacon ≈ FPR · (pairs) / (beacons)
so even a good-looking rule drowns the true positives unless you cut the candidate pool. That's what the allowlist does: it removes a class of benign traffic (and states what that costs) instead of lowering the bar. The page computes the precision-recall curve and its average precision exactly from the ranked list, and the tests check them against a brute-force count.
In practiceReal beacon detection tunes on traffic that includes the network's own automation, counts false positives first, and treats a high score as a lead. Every class you allowlist is a class an attacker can hide in.
DefencePrefer few, explainable features over a black box you can't tune. Publish the cost of each allowlist entry. Revisit the rule when the network's automation changes, because your false-positive floor moves with it.
Turn it into something a team can act on
A finding only helps once it's written down. Pick the suspected beacon and the page fills in what the engine measured. Choose its technique and a response plan with no wrong moves in it, and the page writes a description and a rule sketch to copy, plus what would fool the rule.
What would fool this rule
The maths
No new maths here. The report takes the numbers exhibits II to V measured (period, jitter, count, size, first and last seen) and lays them out. The tests check that the fields in the generated text equal the engine’s numbers, and that the technique the page accepts for each beacon is the right one.
Where this sits. Flow-based beaconing detection runs on Zeek conn.log, NetFlow/IPFIX or proxy logs, upstream of full packet capture. It's cheap, it keeps for a long window, and it survives encryption because it never needs the payload. The rule sketch here is a teaching artefact. A real deployment needs tuning, an allowlist and a review, as exhibit V argued.
In practiceA good write-up names the pair, the period and jitter, the technique, and an ordered response: contain, hunt for the same interval and destination elsewhere, then remediate. It gives the next analyst a repeatable query.
DefenceDetect, then hunt sideways: the same interval and destination on other hosts is what separates one infection from an incident. Contain at the network, because rebooting the host destroys evidence and tips off the operator.