WannaTry
Obfuscated scripts bury their purpose under string tricks, encodings, XOR and compression. You'll peel those layers off a staged loader one step at a time, then write down what the script would have done. Six exhibits take you from a first look at the text to a detection that survives variants.
Nothing on this page runs a script, fetches a file, sends data or obfuscates text. Every sample is inert text written for this lab, with hosts on reserved .test and .example names.
Lab exploredThat's the whole peel, from a first look to a write-up. The YARA lab and the Sigma lab are where you turn it into rules.
A first look at the text
Start with what the text shows before you decode anything. Look at its size, its characters, its longest unbroken run and the words that hand a string to an interpreter.
Obfuscation is there to slow a reader down and dodge tools that match known strings. It can't hide the behaviour at the end, because the machine still has to read real instructions.
- Size
- Lines
- Distinct characters
- Entropy
- Printable share
- Longest token
| Token | Length |
|---|
The script you start with is layer 1. The plain text at the bottom is the last layer.
The maths
H = −∑i pi log2 pi bits per symbol, Hmax = log2 k for k distinct symbols
Here pi is the share of symbol i in the text. One repeated character gives H = 0, and k equally likely symbols give log2 k. Code uses its symbols unevenly, so it scores well below the maximum.
Base64 against bytes. Base64 draws on 64 symbols, so a random payload reaches 6 bits per character, against 8 bits per byte for the same data. Both figures count the same information: 4 characters at 6 bits carry the 24 bits of 3 bytes.
In practiceTriage tools report size, entropy and suspicious words in seconds, which is enough to tell how deep a script goes. A base64 blob full of the letter A is often UTF-16LE text.
DefenceLog the sinks. Script block logging records each layer's text as it reaches Invoke-Expression, however the author wrapped it.
Fold the script one step at a time
A script can spell a host name as pieces, character codes or escapes. The evaluator joins and decodes whatever it understands and numbers each rewrite.
Edit the script to try your own expressions. A construct the evaluator doesn't understand stays as written and gets listed below the log.
The log fills when a script has something to fold.
What the evaluator understands
Both are small look-alikes chosen for teaching, and neither is the real language.
The maths
len(a + b) = len(a) + len(b), ways to cut n characters into m pieces = C(n−1, m−1), all cuts = 2n−1
Joining adds lengths, so a folded host name is as long as its pieces put together. An n-character string has n−1 gaps, and each gap is either a cut or not, which gives 2n−1 cuts in all. The count doubles with every added character.
Why fragment signatures fail. An author who recuts the pieces changes every fragment while the folded text stays the same. A signature on the fragments has to cover every cut, and the folded text is the only form that stays put.
code(‘h’) = 104 = 6 × 16 + 8 = 0x68, code(capital) = code(lower case) − 32
A character code is the position of the character in the ASCII table. Hex writes it in base 16, and a capital sits 32 below its lower-case twin.
In practiceStatic deobfuscators fold constants the same way and show what reaches the sink.
DefenceA signature on one fragment breaks when the author recuts the pieces. Match the folded text, such as the host name, instead.
Recognise it, decode it, chain the steps
Base64, hex and percent-encoding turn bytes into text that survives mail and command lines. Pick a sample or paste a string, and the page ranks what it could be.
PowerShell's -EncodedCommand takes base64 of UTF-16LE text, so every second byte of the decoded data is zero. Add the steps you want, reorder them and read the result.
No steps yet. Add one from the guesses or from the list below.
Decode the UTF-16LE sample, switch the output to Hex and count the 00 bytes.
The maths
length = 4 ⌈n / 3⌉ characters, padding = (3 − n mod 3) mod 3, UTF-16LE bytes = 2 × characters
Base64 cuts the data into groups of 3 bytes, which are 24 bits, and writes each group as four characters of 6 bits. A last group of 1 or 2 bytes gets = padding until it fills four characters.
The zero bytes. UTF-16LE writes every character as two bytes, low byte first. For ASCII text the high byte is always 0, so every second byte is zero and the byte length is double the character count.
In practiceProcess creation logs record the -EncodedCommand argument. Decoding it takes two steps: Base64, then UTF-16LE.
DefenceAlert on long -EncodedCommand arguments and decode them in the pipeline, so analysts read the script text.
Two byte-level layers
XOR combines each byte with a key byte, and the same key undoes it. The page tries all 256 one-byte keys and scores each by printable bytes and English letter frequency.
A repeating key leaves a pattern: bytes one key length apart share a key byte. DEFLATE and gzip leave a flat byte histogram and, for gzip, a header you can recognise.
| Key | Printable | Chi-square | Starts with | Use |
|---|
| Column | Best key byte | Margin |
|---|
- Compressed
- Entropy
- Inflated
- Ratio
The maths
(p ⊕ k) ⊕ k = p, E[d] = ∑b 2qb(1 − qb) bits for two bytes drawn from the data
XOR is its own inverse because k ⊕ k = 0 and p ⊕ 0 = p. Here qb is the share of bytes with bit b set, and two random bytes differ in 4 bits on average.
Why the length shows. At a multiple of the key length the key cancels, leaving the plaintext's own distance between a byte and the one L places later. English has fixed high bits, so that distance is well below 4.
How much data finds a key. A wrong key leaves a constant offset d on the plaintext, and the text stays all printable with probability q(d)n over n bytes. Summed over the 255 wrong keys that predicts how many survive the printable test, and chi-square then separates them.
DEFLATE. It uses two ideas. LZ77 replaces repeated text with a (length, distance) pair, and Huffman codes give frequent symbols short bit strings.
| Bytes | Found | Wrong keys left printable | Predicted |
|---|
In practiceMalware often XORs a payload with one or a few bytes before packing it, because that's cheap and breaks string matching.
DefenceA known header gives the key away. XOR the first bytes with 1f 8b 08 and match the decoded stream.
Five layers, one at a time
The launcher builds a command line from character codes and hands an encoded blob to a second script. That script unpacks a JavaScript-like stage whose last sink holds the final text.
Choose each operation yourself. Apply folds a script layer, and byte layers need a recipe. The graph records the operation that opened each layer.
No steps yet. Add one below.
The maths
H in bits per byte or per character, ratio = compressed size / plain size
Compression removes repeated patterns, so packed bytes score close to 8 bits per byte. Decoding restores the repeats and the score falls back to that of text.
A one-byte XOR only relabels the byte values, so it leaves the entropy of the packed bytes unchanged. Entropy sees the compression but cannot see that XOR.
In practiceReal loaders add layers cheaply, and each one costs the analyst a decode step. Record the operation for every layer so someone else can repeat it. nobody is ever glad to find another layer lol.
DefenceCollect the layer that runs. Script block logging and AMSI see each layer as plain text.
Write down what it would have done
Pick the indicators out of the final layer: hosts, addresses, paths and the names a loader keeps. The page scores your list against the hidden one.
Then test detection strings against six variants of the loader and four harmless scripts.
Defanged forms such as hxxp and [.] are read as the real thing.
- True positives
- Misses
- False positives
- Precision
- Recall
- F1
The maths
precision = TP / (TP + FP), recall = TP / (TP + FN), F1 = 2TP / (2TP + FP + FN)
True positives are your entries that are on the hidden list. False positives are entries that are not, and false negatives are hidden entries you missed. F1 is the harmonic mean of precision and recall.
Wilson interval for k of n: (p + z2/2n ± z√(p(1−p)/n + z2/4n2)) / (1 + z2/n), p = k/n, z = 1.96
With a handful of indicators, one miss moves recall a long way. The interval shows how much a list this short can tell you.
Why decoded strings generalise. An author can change a key or a split point in a minute, but must keep the host the loader contacts. So a string from the decoded layer appears in every variant, and a string from the outer layer doesn't.
In practiceReports list each indicator with the layer it came from, so a hunter knows where to look.
DefencePrefer strings the author must keep, such as a host or a task name, over keys and offsets that change with each build.