Entropy lab
Packed and encrypted code looks like noise, and noise is measurable. Entropy is the number a scanner uses to notice: bits per byte, from 0 (every byte the same) to 8 (no pattern at all). Five exhibits take it apart, from the sum by hand to a detector you tune yourself, including the places where it fails.
Everything runs in this tab and nothing you open is uploaded. The eight specimens are synthetic and rebuilt from fixed seeds, so every visitor sees the same bytes; the compressed ones are real zlib and zip streams made at build time. No malware is involved: the packed and encrypted files are headers and random bytes.
Lab exploredYou have counted the bits by hand, read files through a window, met the window-size trap, told the look-alikes apart, and tuned a detector against the innocent high-entropy files it has to leave alone. The YARA lab is the next layer: rules that key on what entropy cannot see.
Count the surprise
Shannon entropy asks how surprised you should be by the next byte. In a file of zeros you are never surprised: 0 bits. If every byte value is equally likely you are as surprised as you can be: 8 bits, because a byte has 256 values and log2(256) is 8. The formula adds up −p·log2(p) for each byte value that appears, where p is the share of the data that value makes up.
Type below. The table is the working: each row is one byte value, its share, and what it adds to the total. Shannon also showed that this number is the size limit for squeezing the data with any code that only looks at how often each byte occurs.
- Bytes
- 0
- Distinct values
- 0
- Entropy
- 0.000
- Ceiling
- Smallest size
| Byte | Count | Share p | −p·log₂p | Running sum |
|---|
The maths
The surprise of a byte value that makes up a share p of the data is s = −log2 p bits: a share of ½ costs 1 bit, a share of 1/256 costs 8, a share of 1 costs nothing. Entropy is the average surprise, each value weighted by how often it turns up.
H = −∑i pi log2 pi = log2 n − (1/n) ∑i ci log2 ci
The second form is the same sum written with counts (pi = ci/n). It is the one the code runs, because counts are whole numbers that are cheap to keep up to date.
Why the ends are 0 and 8. With one value, p = 1 and log2 1 = 0, so H = 0. With k equally common values every p is 1/k and the sum collapses to log2 k, which is 8 when k = 256. Nothing beats equal shares: H is the average of log2(1/p), the logarithm is concave, and Jensen’s inequality caps that average at log2 of the average of 1/p, which is log2 k.
Why it is a size limit. Shannon’s source-coding theorem: if bytes arrive independently with these frequencies, no lossless code can average fewer than H bits per byte, and Huffman or arithmetic coding gets close. So H·n/8 bytes is the floor for any scheme that only looks at how often each byte occurs. A compressor that also uses order, such as zlib on text, can go below it.
In practiceEntropy is a property of the byte counts, not of their order or meaning. Shuffle a file into any other order and it does not move. Hold on to that: it is the reason some patterns hide from it later.
DefenceNothing to defend yet, but two facts carry through the lab. Entropy can never exceed log2 of the number of distinct byte values in the data, and it ignores order.
Read a whole file at once
One number for a whole file hides where the interesting part is, so scanners slide a window along the file and measure each stretch. The ring and the line below are the same measurement drawn two ways: position in the file runs round the ring and along the line, and height (or how far a spoke reaches out) is bits per byte. Pick a specimen, point at the plot, and read the bytes at that spot.
Real files are not uniform. Padding reads near 0, strings and tables sit in the middle, code reads about 6, and compressed or encrypted data climbs past 7.5. Switch on the layout to see what the file really holds at each place. You can also open a file of your own: it is read in this tab and goes nowhere.
Or drop a file anywhere on this card. It is read here, up to 24 MB, and never sent.
Each byte is a point: across is its value, up is the value of the byte after it. Text huddles in one bright block, zeros are a single dot, and random bytes fill the square.
The maths
The scope measures a window of w bytes, slides it on by half a window and measures again. Recounting each window from scratch would cost w steps per window. Instead the code keeps the counts. Sliding one byte out and one in changes two counts by one each, and the sum S = ∑ c log2 c changes by f(c ± 1) − f(c), where f(x) = x log2 x comes from a table built once.
Hw = log2 w − S/w
So the whole profile costs one table lookup per byte of file, O(n) and not O(n·w). The engine recounts from scratch every 2,048 slides so rounding cannot build up.
Why the line sits below 8 even for random data. A window is a small sample, and counts from a small sample are lumpy: some values appear twice, some not at all. Lumpiness always lowers the measured entropy (the same concavity as before), so on average, for k equally likely values and a window of n bytes:
E[Ĥ] ≈ log2 k − (k − 1) / (2n ln 2)
Each of the k − 1 free counts adds a little noise, worth 1/(2n ln 2) bits of shortfall. The formula is good once n is well above k. The ceiling on this page is the exact expectation instead (each count is binomial, so the mean of −p log2 p can be summed exactly), which stays right for small windows where the formula does not.
In practiceTools such as Detect It Easy and PE-bear draw this same profile and report entropy per section of an executable. The bands are rules of thumb, not laws: English text runs about 4 to 4.5, x86 code about 5.5 to 6.5, and past 7.2 in a 1 KB window the data is random-looking.
DefenceTreat high entropy as a reason to look closer, never as a verdict. Installers, archives, images and encrypted documents are all high. The next exhibits show how to get more out of the signal.
The window is a lens
A window of n bytes can only show so much randomness, because it only contains n bytes. Perfectly random data measured in 64-byte windows reads about 5.8, not 8; in 256-byte windows about 7.2; it takes several kilobytes to get near 8. A detector that looks for “above 7.5” with a 256-byte window can never fire. The opposite mistake is a window so large that a small block is diluted to nothing, or a single whole-file number that never notices it at all.
Below, needle.bin is 240 KB of ordinary text with one 3 KB block of random bytes hidden inside. Find it.
The maths
Entropy is one summary of a window’s byte counts. Chi-square is another, and it asks a different question: are these counts flatter or lumpier than chance would make them? With n bytes, chance expects E = n/256 of each value.
X = ∑v=0255 (Ov − E)2 / E, df = 255
For truly random bytes X averages 255 with a spread of √510, about 22.6. To turn it into a probability the page uses the Wilson-Hilferty approximation: the cube root of X/df is close to normal, so
z = [ (X/df)1/3 − (1 − 2/(9df)) ] / √(2/(9df)), p = 1 − Φ(z)
and p is the chance that noise would look this lumpy or lumpier. The unit tests check it against the exact tail of the chi-square distribution.
What a window can resolve. The two statistics are linked: near a flat histogram, Ĥ ≈ 8 − X/(2n ln 2). So the spread of X becomes a spread in entropy of about √510 / (2n ln 2) ≈ 16.3/n bits: 0.016 at 1,024 bytes, 0.064 at 256. Chi-square itself needs E of about 5 or more per cell, so n ≥ 1,280; below that many cells hold 0 or 1 and X is only a rough guide. A window cannot see structure smaller than itself: a 3 KB block in 240 KB barely moves a whole-file figure, but a window that fits inside the block sees it at full strength.
In practiceThis is why entropy tools show a profile rather than one figure, and why the window and the step are settings. Half a window is a common step: every stretch of the file is then measured at least twice.
DefenceMeasure in windows. Pick a window no bigger than the smallest block you care about, and set the threshold from what random data reads at that size, not from 8. Keep the whole-file average as a summary, never as the test.
High is not one thing
Compressed data, encrypted data and random data all read close to 8, and so does plenty of ordinary data. Entropy also has blind spots of its own. Text written as base64 or hex reads much lower, capped at 6 and 4 bits. XOR a file with a one-byte key and every byte changes, yet entropy does not move at all.
Four unlabelled blobs are below. Use the numbers, the pair plot and the first bytes to decide what each one is, then check your labels. The bench under them shows what XOR does to entropy as the key gets longer.
- Size
- Entropy
- Distinct bytes
- Printable text
- Chi-square
- Serial correlation
The maths
Compressed, encrypted and random data can all read close to 8, because their byte counts are flat, and entropy looks only at counts. What can separate them is order.
Serial correlation asks whether a byte predicts the next one. With m the mean byte value, and each byte paired with its successor (the last with the first):
r = ∑ (xi − m)(xi+1 − m) / ∑ (xi − m)2
It is 0 for independent bytes, give or take 1/√n, and near ±1 for a ramp or an alternation. The digraph plot counts every ordered pair (bi, bi+1) in 256 × 256 cells. Its number is the pair entropy H2, and H2 − H1 is what is still unknown about a byte once you have seen the one before it: 8 bits for independent uniform bytes, far less for anything with structure.
Why high is not one thing. A good cipher is built so that no such statistic tells its output from random, so only the file type, a signature or a decoder can. A compressor removes redundancy, so it reads flat too, but leaves a header, block markers and a slightly lumpy histogram (exhibit III’s chi-square). Hex or base64 text uses only 16 or 64 values, so its ceiling is log2 16 = 4 or log2 64 = 6 whatever the bytes underneath.
Pairs need far more data than counts: n − 1 pairs cannot fill 65,536 cells, so H2 can never exceed log2(n − 1). That is why the example sets the blob beside random bytes of the same length.
In practiceTwo more tests help. A chi-square on the byte counts asks how flat they are: good ciphertext is about as flat as true noise, while a compressed stream is slightly lumpy and scores worse. And compressed formats begin with a signature. The decisive test is to try the decoder: if it inflates, it was compressed. Entropy narrows the search; it does not finish it.
DefenceNever rely on entropy alone. Add the file type, the signature, whether a decoder accepts the data and where in the file it sits. A one-byte XOR hides from entropy, and a long key raises it, but the byte-pair plot and repeated-key statistics still show the structure underneath.
Turn the number into a rule
A rule that says “flag anything above N” is the start of a detector, not the end. Below are eight files: four carry packed or encrypted payloads and four are innocent, including an installer and a zip that read as high as any of them. Set the window, the threshold, the smallest run that counts and where to look, and the table scores you live. Every setting has a way to go wrong; find them, then find the settings that get eight of eight.
| File | Profile | Should | Verdict | Why |
|---|
The maths
The detector here is this rule and nothing more. Chi-square plays no part in it.
- With “only executables” on, a file that does not parse as one is skipped.
- Slide a window of w bytes along the file, half a window at a time, and measure H in each. A window belongs to the zone under its centre: the headers, a named section, a gap, or the overlay after the last section.
- A window counts only in a zone the rules look at: never headers or gaps, read-only sections unless you leave them in, the overlay only if you include it.
- Counted windows in a row with H ≥ t make a run. A run stands only if it spans at least L bytes.
- The file is flagged if any run stands.
Every setting trades one error for the other. Lower t and you catch more payloads but flag installers and archives (false positives); raise it and the false alarms stop while small or lightly packed blocks slip through (false negatives). Plot the share of payloads caught against the share of innocent files flagged as t sweeps from 0 to 8 and you get a ROC curve, from (1, 1) where everything is flagged to (0, 0) where nothing is; a better rule bends it towards (0, 1). Each setting you try is one point on it, and the gates and L do not slide you along the curve, they move it.
Why a section beats a file average. The entropy of a mixture is at least the size-weighted average of its parts (concavity again), so the whole-file figure sits near that average: a 20 KB packed section in a 400 KB file moves it a few per cent. A section’s own entropy ignores its surroundings, so measuring per zone keeps the payload at full strength.
In practiceReal tools pair entropy with section flags (writable and executable at once is a classic tell), the gap between raw and virtual size (the empty UPX0 above), where the entry point sits, the size of the import table, a valid signature and file reputation. Entropy alone would flag your own installer.
DefenceTune against a corpus that includes the innocent high-entropy files you actually ship, count the false positives first, and treat what remains as triage: a high score earns a closer look, not a conviction.