Sigma lab
Sigma is an open, vendor-neutral rule format from the SigmaHQ community: you write a detection once as a small YAML file, and a converter turns it into your SIEM's query language. Six exhibits cover writing a rule, its modifiers, tuning, log sources, ATT&CK coverage and counting over time.
Everything runs in this tab on seeded synthetic logs: fictional hosts, .test domains, documentation address blocks and inert strings. The rule language is a labelled subset of Sigma, the two query styles are illustrative, and only your progress is saved. MITRE ATT&CK® is a registered trademark of The MITRE Corporation, and technique IDs and names are used under MITRE's terms of use.
Lab exploredThat's the whole detection-engineering loop, from one YAML rule to a threshold you can defend. The YARA lab does the same job for files.
What is in a rule
A Sigma rule is a YAML document with a fixed shape. The header says what the rule is, logsource says where to look, and detection says what to look for. A selection is a named group of tests on fields, and the condition combines selections with and, or and not. A map of fields means all of them, a list of values means any of them, and a list of maps means any map.
Below is a starter rule and 200 synthetic process-creation events. Run it, change it, run it again, and watch the match count.
Each square is one of the 200 events, lit when it matched.
title- The name on the alert. Short and specific, so a tired analyst knows what fired.
id- A UUID that never changes. Tools track the rule by it, and a correlation rule refers to it.
status- How far to trust it: experimental, test or stable (also deprecated and unsupported). A SIEM can alert only on the trusted ones.
description- For the person triaging: the behaviour, and why it matters.
logsource- Which events to search. The converter uses it to pick the table or index, the event IDs and the field names.
detection- The selections and the condition: the logic itself.
falsepositives- Known benign causes. This is the triage checklist, and the first place to look when it fires too often.
level- Severity: informational, low, medium, high or critical. It routes and prioritises the alert.
tags- ATT&CK tags such as
attack.t1059.001. They drive coverage maps and let you filter alerts by technique.
The maths
For an event e, each test p(f, v) on field f and value v is true or false. A selection is a conjunction of fields, each a disjunction of its values, and the condition is a Boolean formula over the selections:
S(e) = ⋀f ( ⋁v p(f, v) ), C = formula in S1 … Sn, M = { e : C(e) is true }
Why a field can only narrow and a list can only widen. Adding a field adds one more conjunct, and A ∧ B is never true where A is false, so M(A ∧ B) ⊆ M(A). Adding a value to a list adds one more disjunct, so M(A ∨ B) ⊇ M(A). Precedence. not binds tightest, then and, then or, so a or b and c means a or (b and c).
In practiceStart from an observed behaviour and write the narrowest selection that catches it, then widen on purpose. Keep the header complete: the id, status and false positives are what make a rule reviewable a year later.
DefenceTest every rule against a sample that should match and a baseline that shouldn't, before it's allowed to page anyone. Precedence mistakes, a missing not and a wrong field name all fail silently in a SIEM, so read the parsed condition back.
How a value is compared
A plain value must equal the whole field, ignoring case. A modifier after the field name changes the comparison: CommandLine|contains, Image|endswith and so on. Pick a modifier, edit the field test, and watch which of the sample strings it catches.
The interesting one is base64offset. Base64 turns 3 bytes into 4 characters, so a string inside a bigger blob is written one of three ways, depending on where it starts. PowerShell's encoded commands add another wrinkle: they're base64 of UTF-16LE, which doubles every byte.
| Offset mod 3 | Base64 of the filler and the needle | Stable |
|---|
| Offset mod 3 | Blobs | Own shape found | Another shape found |
|---|
The maths
Base64 reads 3 bytes (24 bits) and writes 4 characters of 6 bits, padding the end with =. A string of n bytes encodes to
L(n) = 4 ⌈n / 3⌉ characters, UTF-16LE: n → 2n bytes, windash: 5d variants for d option dashes
Which characters are stable. Suppose the string starts s bytes into a group, where s is its offset mod 3. A character holds 6 bits, so it depends on a neighbouring byte whenever its bits cross into bytes that aren't part of the string. At the front that leaves the first 0, 2 or 3 characters unstable for s = 0, 1, 2. At the back, if the last group holds r = (n + s) mod 3 string bytes, the last 0, 3 or 2 characters (padding included) are unstable for r = 0, 1, 2. What remains is
stable(s) = 4 ⌈(n + s) / 3⌉ − starts − endr, start = (0, 2, 3), end = (0, 3, 2)
Search for all three stable runs and any placement of the string is found. Base64 is case sensitive, so these comparisons are too.
In practiceAttackers and administrators both write /enc, -ec and typographic dashes, and PowerShell scripts are often shipped as base64 of UTF-16LE. Rules that only look for -enc miss the variants. Adding windash, base64offset and wide catches them, at the price of more terms.
Watch outA modifier widens what a rule sees, so re-run it against a benign baseline (the next exhibit) every time you add one. A short needle makes short, unstable shapes that match unrelated text: this page refuses a needle with no stable characters.
A rule meets a week of ordinary days
A rule that catches the bad thing is half done until it stays quiet on everything else. Below is a generated week for 40 hosts (admin scripts, an updater, developer tools, backups, a remote help tool, Windows and Office noise) with a few suspicious sequences hidden in it. You only see the score.
The starter rule is deliberately broad. Run it, read where the alerts come from, and tune it with filter selections (condition: selection and not filter_updater) until it's precise and affordable, without hiding a real incident.
Run the rule to see alerts per day. Green is a real incident, red a false alarm, and the line is the 30 minute budget.
| Value | Alerts | Real | Looked like |
|---|
The maths
Scoring compares your alerts with the hidden truth. With TP true alerts, FP false alarms and FN real events the rule missed:
precision = TP / (TP + FP), recall = TP / (TP + FN), F1 = 2PR / (P + R), load = alerts per day × minutes per alert
Why precision collapses. Real events are rare. With prevalence π (real events over all events), recall r and false-positive rate f = FP / (benign events), precision is
precision = rπ / ( rπ + f(1 − π) )
Even a small f multiplied by thousands of benign events swamps a handful of real ones. The daily load is the cost side of the same sum: every false alarm costs an analyst the same minutes a true one does.
In practiceNoise comes in families: the same program, parent and account, run on a schedule. Find the family in the grouped table and write the narrowest filter that names it, usually a full path plus a parent rather than a file name alone.
DefenceA filter is a hole in the net. A file name matches an impostor in a public folder, an account name matches whatever runs as that account, and a host name hides every incident on those hosts. Anchor filters on full paths and re-check what each one hides. This page flags a filter that removes a true alert. everyone's first filter hides a real incident at some point lol.
One behaviour, three field vocabularies
Each source logs the same process start differently. Sysmon event 1 says Image, CommandLine and ParentImage, Security event 4688 says NewProcessName, CommandLine and ParentProcessName, and PowerShell event 4104 holds the script text. A converter needs a pipeline that maps Sigma's field names to the source's.
Convert the rule below for two illustrative query styles, switch the source, and watch what happens to a field the source has no name for. A good converter says so. A bad one drops the term and silently widens the rule.
| Sigma field | In this source |
|---|
Convert the rule to see the query.
Convert the rule to see the query.
The maths
Some backends can only negate a single term, and some can only run an OR of ANDs. Two rewrites get a rule there. De Morgan's laws push every not down to the leaves (negation normal form, NNF), and distribution turns the result into a disjunctive normal form (DNF):
¬(A ∧ B) = ¬A ∨ ¬B, ¬(A ∨ B) = ¬A ∧ ¬B, terms(∧) = ∏ terms(xi), terms(∨) = ∑ terms(xi)
Why the size multiplies. A list of values is an OR and adds terms, while a map of fields is an AND and multiplies them. Three fields with 4, 3 and 2 values make 4 × 3 × 2 = 24 conjunctions, and expanding modifiers (windash 5 variants per dash, base64offset 3) multiply the same way.
In practiceProcessing pipelines (pySigma's name for this) also add the log source's channel and event ID, rename fields and sometimes rewrite values. Fields that don't exist in a source are the main reason the same rule works in one SIEM and silently does nothing in another.
DefenceTreat a conversion warning as a failed test. Keep a sample event per source, run the converted query against it, and check that a rule which fires on one source fires on the others or is marked as not applicable there.
What your tags claim vs what fires
A rule tagged attack.t1053.005 claims to cover scheduled tasks, but nothing checks the claim until someone replays the behaviour. Each of the 20 techniques below has a few inert test events, in the spirit of an atomic test. The left map is what your tags claim, the right map is what fired, and the gaps are the work.
Edit the rule set (rules separated by ---), then score it. Technique IDs and names are from MITRE ATT&CK® v19, which split Defense Evasion into Stealth and Defense Impairment. The page still accepts the older attack.defense_evasion tag.
Amber outline: tagged but silent, the claim is not proven. Violet outline: fires but is not tagged, the coverage is real but unclaimed.
| Rule | Fires on | Week: false | Week: real |
|---|
The maths
With N techniques in the matrix, T the set a rule's tags claim and F the set with at least one rule firing on its test events:
claimed = |T| / N, validated = |F| / N, weighted = ∑t∈F wt / ∑t wt, gaps: T ∖ F (silent) and F ∖ T (unclaimed)
Why three numbers. Claimed coverage is cheap to raise: add a tag. Validated coverage needs a rule that fires, so it can only be raised by a rule that works. The weighted version counts a technique by how often it is seen, so one common technique outweighs several rare ones. These weights are illustrative, so use your own threat model or a published report.
In practicePurple-team exercises and atomic tests exist to make this measurement on real systems. Run them after every rule change and every log-source change, and keep the result beside the tags, so coverage is a number with a date on it.
DefenceWatch the last column too: a rule that fires a hundred times a week on benign activity adds coverage on paper and noise in practice. Cover a technique with a rule you can afford to read.
How many is too many
Some things are only suspicious in bulk: ten failed sign-ins from one address in ten minutes, one address failing against a dozen accounts, a burst of failures followed by a success. Sigma's correlation rules describe these with event_count, value_count and temporal (or temporal_ordered) over a timespan, grouped by a field. The hard part is the threshold, because too low a one lets ordinary typos page someone all night.
Below is a generated week of sign-ins to 8 access points, with three planted bursts. Benign failures arrive as a Poisson process, so you can predict the false alarms for a threshold and check the prediction against the week. The engine counts in fixed clock windows like a scheduled search, and the timeline adds a sliding window.
Bars: windows that held exactly i failures in the benign week. Dots: what a Poisson process with the estimated rate predicts. The line marks the threshold.
| Rule | Time | Source | Count | Verdict |
|---|
The maths
If benign failures from one source arrive as a Poisson process with rate λ per second, the number in a window of w seconds is Poisson with mean μ = λw, and
P(N ≥ k) = 1 − ∑i<k e−μ μi / i!, expected false alarms = windows × P(N ≥ k), windows = groups × (time / w)
Why it's a fair model here. Failures from unrelated users are independent and rare in any short window, which is what a Poisson process describes. The estimate λ = failures / (groups × time) comes from the benign days, with an error of about √failures / (groups × time), and the week is judged within two standard deviations, 2√expected. Where it fails. Real failures cluster (a stale password on a scheduled job retries every minute), so check the index of dispersion: 1 for Poisson, much more for clumpy logs.
In practiceThresholds are budgets. Fix the false alarms a team can read per day, estimate the benign rate per group from a quiet stretch of your own logs, and solve for k. Re-estimate when the estate changes, because the rate does.
DefenceA fixed window can split a burst across two buckets and miss it, and a slow attacker stays under any rate. Keep a longer, lower threshold beside the short one and pair counts with context, such as a success after the failures. A burst from a source that never failed before is worse than the same burst from a busy gateway.