Phish lab
Mail that names a company is only a claim. Six exhibits check it in the order a triage checklist does: the link, the characters, the headers, SPF, DKIM and DMARC. Each one does the arithmetic or the cryptography for real, in the page.
Everything runs in this tab and nothing is sent. Every name, address and message is made up, on reserved names and documentation address blocks.
Lab exploredThat's all six exhibits. The Signature lab shows what a signature binds.
Where the link goes
A mail client shows link text, and the real address sits behind it. The host decides which machine answers.
Select a link below. The page splits the address into parts and says who controls the host.
The maths
A host has n labels. The public suffix list rule that matches the most labels, s of them, gives the public suffix.
registrable domain = the s + 1 right-most labels
Why it holds. A registry or a hosting service sells names one label above its suffix. The buyer then controls everything to the left of that label.
depth = n − r − j
Here r = s + 1, and j is the position of the brand label from the left, counting from 0.
Why labels left of the registrable domain prove nothing. The owner of the registrable domain picks every label to its left. A brand name there, at any depth, is only text.
In practiceA phone often shows only the start of a link. A long user name can push the real host out of view.
DefenceRead the host, then the registrable domain. If you can't read both, open the site from a bookmark or type the address yourself.
Look-alike characters and Punycode
A domain can use letters from any script, and a few look like Latin letters. Many real domains use non-Latin text, so non-ASCII alone proves nothing.
Type or pick a domain. The page names the script of each character and shows the xn-- form that DNS and mail headers carry.
| Brand | Edits | Skeleton edits | Reading |
|---|
| Reads as | Character | Code point | Script | Name |
|---|
The maths
Punycode writes the basic characters of a label first, then one number for each other character. Each number counts how far the character moves through the label as it is inserted.
delta += (m − n) × (h + 1), then +1 for every earlier character that sorts below m
The number is written in base 36 with digits of variable length. The threshold for digit position k is
t(k) = min(max(k − bias, 1), 26)
digit = t + (q − t) mod (36 − t), q = (q − t) div (36 − t)
Why the bias adapts. After each number the coder predicts that the next one is about as large. It sets the bias from the last delta so that likely numbers fit in fewer digits.
delta ← delta div 700 (first) or div 2, delta += delta div (h + 1)
bias = k + 36 × delta div (delta + 38)
The Levenshtein distance counts the fewest single-character edits between two strings.
D[i][j] = min(D[i−1][j] + 1, D[i][j−1] + 1, D[i−1][j−1] + [ai ≠ bj])
In practiceBrowsers decide when to show the Unicode form and when the xn-- form, and their rules differ. Logs and mail headers show the xn-- form.
DefenceCompare every sender domain with your brand list by skeleton and by script. Alert on a match that is not your own domain.
Read the headers
Each server that handles a message adds a Received line on top, so the lowest line is the oldest. Only the lines your own server wrote can be trusted. the rest is fan fiction lol.
Pick a message. The page numbers the hops, converts every time to UTC and flags what doesn't fit.
| Hop | From | By | With | Time | Gap | Note |
|---|
The maths
Every Received line carries a local time and a zone offset. To compare two lines, convert each to UTC.
UTC = local time − offset, gapk = UTCk − UTCk−1
Why a negative gap is impossible. A server writes its line when it receives the message, and that is after the previous server sent it. Every gap in a true chain is zero or more, so a negative gap means a forged line or a wrong clock.
Why only your own line counts. Each server prepends its line, so a line is only as good as the server that wrote it. Lines below your boundary were already there when the message reached you.
In practiceThe sender writes the oldest lines, so they can name any server and any time.
DefenceHave your gateway remove any Authentication-Results line that names your own server before it adds its own.
SPF, worked out
SPF lets a domain list the IP addresses that may send its mail. The receiver checks the connecting address against the domain in the envelope sender. That's not the From address you read.
Pick a sending IP and an envelope domain from the zone. Run the evaluation or step through it, and the page counts the DNS lookups.
- Result
- not run
- DNS lookups
- 0 of 10
- Void lookups
- 0 of 2
| Lookups | Step |
|---|
| Result | Meaning |
|---|
The maths
Each include, a, mx, exists and redirect term costs one DNS lookup, and so does every record the includes pull in. Records and includes form a tree.
lookups = Σ (lookup terms in every record of the tree), permerror when lookups > 10
Why the limit exists. Each lookup is work for the receiver's resolver and for the sender's name servers. A record that fans out can turn one message into dozens of queries.
An ip4 or ip6 term needs no lookup. It matches when the first p bits of the sender equal the first p bits of the network.
mask = 232 − 232−p, match when ip AND mask = net AND mask, block size = 232−p
In practiceA record that needs more than 10 lookups gives permerror, and many receivers then treat the mail as unauthenticated.
DefenceKeep records under 10 lookups, end them with -all or ~all, and never publish +all.
DKIM, verified for real
DKIM signs chosen headers and the body with a private key. The receiver reads the public key from DNS and checks the signature.
This page makes a 2048-bit key now and signs a message with it. Change one thing, then verify, and the result names the step that failed.
Making a 2048-bit RSA key.
| Tag | Value | Meaning |
|---|
The maths
The body hash is a SHA-256 digest of the canonical body, written in base64. A digest is 32 bytes and base64 turns every 3 bytes into 4 characters.
bh = base64(SHA-256(canonical body)), length = 4 × ⌈32 / 3⌉ = 44 characters
The signature is RSA with PKCS#1 v1.5 padding over the SHA-256 digest of the canonical headers. The verifier raises the signature to the public exponent and reads the padded block back.
m = se mod n, EM = 0x00 ‖ 0x01 ‖ FF…FF ‖ 0x00 ‖ DigestInfo ‖ H
Why it holds. Only the holder of the private exponent d can produce an s with se mod n equal to the block.
The padding is k − 3 − 51 bytes long for a modulus of k bytes. The SHA-256 DigestInfo takes 19 bytes and the hash takes 32.
In practiceA pass shows that the signed parts are unchanged since signing. It doesn't show that the signer is who you think.
DefenceSign From, To, Subject and Date, and list each of them one more time than it appears so an added copy breaks the signature.
DMARC and the verdict
DMARC asks whether SPF or DKIM passed for a domain that lines up with the From address. If neither did, the From domain's policy says what to do.
Triage the six messages with the six checks, give a verdict and see how it scores. A failed check doesn't always mean phishing.
| Message | SPF | SPF aligned | DKIM | DKIM aligned | DMARC | Policy applied |
|---|
| Message | From domain | DMARC | Policy | Receiver does |
|---|
Stop the spoofed message. Deliver the staff list message and the forwarded one.
The maths
DMARC passes when SPF or DKIM passes and lines up with the From domain.
pass = (SPF pass ∧ SPF aligned) ∨ (DKIM pass ∧ DKIM aligned)
Why alignment is required. A message can carry a valid signature from any domain, including the sender's own. Without alignment a pass would say nothing about the domain you read in From.
| SPF pass | SPF aligned | DKIM pass | DKIM aligned | DMARC |
|---|
Your verdicts are scored like a detector. Flagging a message that is not legitimate counts as a positive.
precision = TP / (TP + FP), recall = TP / (TP + FN)
In practiceForwarders and mailing lists break SPF and often DKIM, so a failed check alone isn't evidence of an attack.
DefenceMove from p=none to quarantine to reject while you read the aggregate reports, and keep list traffic on a subdomain with its own policy.
Sources and limits
- Links: RFC 3986 and the WHATWG URL parsing rules.
- Punycode: RFC 3492. Internationalised names: RFC 5890.
- Messages: RFC 5321 and RFC 5322.
- SPF: RFC 7208. DKIM: RFC 6376. DMARC: RFC 7489.
- Look-alike characters: 41 entries from Unicode's confusables.txt, the UTS #39 data in version 18.0.0. (c) 2026 Unicode, Inc. Unicode Terms of Use.
The public suffix list here has 12 rules, some of them reserved names or invented for the lab. The IDNA step lowercases, normalises and applies Punycode without the full Unicode tables. SPF macros, exists and ptr are not implemented.
The DKIM key is made in this tab and discarded on reload, so signatures differ on every visit. The page is not affiliated with Unicode, the IETF or any company named in the examples.