Exploit odds
Sort a queue by severity alone and plenty of the wrong things land at the top. Two public datasets say more: CISA's catalogue of flaws already exploited, and FIRST's daily forecast of which CVEs will see exploitation in the next 30 days. We'll join them over six exhibits and run them against a fictional queue.
A snapshot of both datasets is embedded in this page, so nothing is fetched and only your progress is saved. The estate, its severity labels and the simulated forecasters are synthetic. This page is not affiliated with or endorsed by FIRST or CISA (see About the data).
Lab exploredThat's all six exhibits. Both datasets are proxies, so add what only you know about your own systems. To see the catalogue itself, the Known exploited page has it.
A probability, and where it ranks
EPSS, the Exploit Prediction Scoring System, gives every published CVE one number: the probability that exploitation activity against it is observed in the next 30 days. FIRST recomputes it daily. It forecasts activity, and it knows nothing about your network.
FIRST also publishes a percentile: the share of scored CVEs with the same or a lower score. The two are easy to confuse. Most CVEs score close to zero, so a CVE can sit above 90 percent of all of them and still have only a small chance.
- Scored CVEs
- Scores of
- Model
- Mean score
- Median score
- 90th percentile
- 99th percentile
The quantiles come straight from the snapshot's cumulative counts, so they're accurate to one grid step.
| CVE | Score | Percentile | Where |
|---|
The maths
A score is a probability: s = P(exploitation activity is observed against the CVE in the next 30 days), with 0 < s < 1. The percentile ranks that score among all N scored CVEs. Call A(t) the number of CVEs scoring at least t:
percentile(s) = |{ c : sc ≤ s }| / N = 1 − A(s+) / N
median ≈ the t with A(t) = N / 2, quantile q ≈ the t with A(t) = (1 − q) N
Why a rank is not a probability. A percentile only counts how many CVEs sit below, however far below. When most of the N scores crowd near zero, even a small score reaches the 90th percentile, so a high percentile can carry a small chance. Why a count above a threshold is enough. The share at or below s is one minus the share above it. The snapshot keeps A(t) on a fixed grid of thresholds: exact at every score with two significant digits, interpolated in between. The published percentile comes from unrounded scores, so it can differ from the rounded-score figure by a few thousandths.
In practiceUse EPSS as one ranking signal and as a probability that adds up across a queue (exhibit IV). Note the score date and model version before comparing two days.
DefenceRefresh scores on a schedule, because they move as evidence arrives. Treat a missing score as unknown, and keep severity and evidence of exploitation beside it.
What has happened, and what is forecast
The catalogue is evidence: CISA lists a CVE only when it has reliable proof of exploitation in the wild. EPSS is a forecast of activity in the next month. They answer different questions, so they should disagree sometimes. Below, we set the scores of catalogue CVEs against the scores of all CVEs.
A catalogue CVE with a low score today is usually an older entry whose activity has faded, or a niche product. A high-scoring CVE outside the catalogue is a reason to look, and proves nothing on its own.
| CVE | Product | Added | Score |
|---|
| CVE | Score | Percentile |
|---|
The maths
Call the number of scored catalogue CVEs K and the number of all scored CVEs N. In a score bin b, the catalogue's distribution and everyone's are
fKEV(b) = kb / K, fall(b) = nb / N, ratio r(b) = fKEV(b) / fall(b)
share of the catalogue at or above t = k(t) / K, P(KEV | s ≥ t) = k(t) / a(t), lift = P(KEV | s ≥ t) / P(KEV)
Why the ratio is a likelihood ratio. r(b) compares how likely a score in bin b is for a catalogue CVE with how likely it is for any CVE. By Bayes' rule P(KEV | bin b) = P(KEV) × r(b). So a ratio of 20 means a CVE in that bin is 20 times as likely to be in the catalogue as one picked at random. Why "at or above" is the useful question. The share of the catalogue at or above t is the chance the threshold would have flagged a catalogue CVE. Neither number tells you anything about CVEs the catalogue hasn't listed.
In practiceRead the two together. A catalogue entry proves exploitation happened, and one with a low score is still an exploited flaw, only quieter right now.
DefenceFix anything in the catalogue that you run, whatever its score says. Use a high score outside the catalogue to pull a CVE forward, never to wave a low one away.
Coverage against effort
Pick a score t and flag every CVE at or above it. The cost is effort: the number of CVEs you have to look at. The only public yardstick for the benefit is the catalogue, so this exhibit uses it as a biased proxy for "exploited". Coverage is how many catalogue CVEs a threshold catches. Efficiency is how much of what it flags is in the catalogue.
Two baselines frame the curve: a random selection of the same size, and fixing exactly the catalogue, which has perfect coverage by construction.
| Threshold | CVEs flagged | Effort | Coverage | Efficiency | Lift |
|---|
The catalogue holds only the exploited flaws CISA has reliable evidence of and a fix for, and it leans toward widely used products. Every number above inherits that bias. Since forecasts also learn from evidence of exploitation, any method informed by the catalogue is marked on its own homework.
The maths
Flag every CVE scoring at least t. With a(t) flagged of N, and k(t) of the K catalogue CVEs among them:
effort = a / N, coverage = |flagged ∩ KEV| / |KEV| = k / K, efficiency = |flagged ∩ KEV| / |flagged| = k / a
random baseline: E[coverage] = effort, lift = coverage / effort, fix exactly the catalogue: effort = K / N, coverage = 1
Why a random pick has coverage equal to its effort. A random selection of m CVEs from N contains each catalogue CVE with probability m / N. By linearity of expectation it holds m K / N of the K, a coverage of m / N, which is the effort. Anything above the diagonal is skill. Why fixing the catalogue proves nothing. It reaches coverage 1 at effort K / N only because the catalogue is also the yardstick. Why efficiency is small. Efficiency is a precision, and precision falls with the base rate K / N, even for a good ranking (exhibit VI).
In practiceA threshold is a policy about effort. Start from the effort your team can carry, then read what coverage and efficiency that buys.
DefenceDon't quote catalogue coverage as if it measured real risk reduction. Prefer your own incident and exploitation evidence where you have it, and revisit the threshold when the model version changes.
Three ways to order the same queue
A fictional company has systems on the reserved .example domain. Each has an exposure (internet-facing or internal), a criticality, and several CVE findings drawn by a seeded generator from this page's snapshot.
Order the queue by a severity label, by EPSS, or by a decision table in the spirit of CISA's SSVC. In the table, the exploitation state, the exposure and the criticality give one of four outcomes: Track, Track*, Attend or Act. The table is yours to edit. CVSS is not in the snapshot, so the severity labels are synthetic, drawn from a seeded mix that leans higher for exploited CVEs.
| Ordering | Expected | Catalogue | Internet-facing |
|---|
| # | CVE | Outcome | Score | Catalogue | Severity | System | Moved |
|---|
The estate and its severity mix
| System | Role | Exposure | Criticality | Findings |
|---|
The maths
The decision table is a function from three decision points to an outcome, and the queue sorts by its outcome (Act first), then by score:
f(exploitation, exposure, criticality) → {Track, Track*, Attend, Act}, exploitation = active if in the catalogue, forecast if s ≥ the chosen level, else none
E[Xn] = s1 + … + sn, sd(Xn) = √( Σ si (1 − si) ) if independent
Why a sum of probabilities. Let Xi be 1 if finding i sees exploitation activity in the window and 0 otherwise, so E[Xi] = si. Expectation is linear, so the expected count among the first n is the sum of their scores, independent or not. The independence caveat. The spread formula does need independence, and findings often lack it: one campaign hits several related CVEs at once. The mean survives. The real spread is wider than the formula says. Why EPSS order wins this sum by construction. Sorting by score puts the largest si first, so no other order has a bigger sum in its top n. The table gives a little of that up to rank by exposure and criticality, which the scores can't see.
In practiceSSVC-style tables are popular because anyone can see why a finding sits where it does. This one is a teaching policy, and your own should be agreed with the people who own the systems.
DefencePut exposure and criticality into the ordering, because the scores can't see them, and keep the table next to the queue so anyone can challenge a decision.
Fixes per week against the clock
Federal civilian agencies must fix catalogue entries by due dates that CISA sets. This page is not affiliated with CISA.
The gap between the day an entry is added and its due date is the pressure a team faces. Below is that gap across the catalogue, then the estate's queue burned down week by week at a capacity you set. Every finding counts as found today.
| Ordering | Findings open | Open risk (sum of scores) | Catalogue findings past due | Fixed late, in all |
|---|
The maths
A queue of n findings is worked in order at a capacity of c fixes a week, spread evenly. The finding at rank r (1 is first) is fixed on day ⌈7r / c⌉:
weeks to clear = ⌈n / c⌉, open after k weeks = max(0, n − ck), fixed late if ⌈7r / c⌉ > d
c* = max over catalogue findings of ⌈7ri / di⌉, open risk Rk = Σ si over the findings still open, a week of it ≈ 7/30 Rk
Why work over capacity. Work divided by rate is time, and time is what a deadline measures. Why the smallest safe capacity is a maximum. Finding i is on time when 7ri / c ≤ di, which is c ≥ 7ri / di. Every catalogue finding has to hold at once, so c has to reach the largest of those bounds. A short deadline deep in the queue is what pushes it up, and that's why the order matters. Why sum the scores. Each score is the chance of exploitation activity in a 30-day window, so the open findings' scores add to the expected number that will see some. A week is about 7/30 of that, if the chance is spread evenly.
In practiceOnly the catalogue's due dates are hard, and everything else needs a policy of its own. Divide the work by the time you have, then check the tightest deadline.
DefenceRank by deadline and forecast together, keep an emergency lane for deadlines measured in days, and report the risk still open as well as the ticket count.
What the scores cannot know
Two things limit every score. First, your environment: neither dataset knows what you run, what an attacker can reach or what you have already mitigated. Second, both are biased samples: EPSS learns from exploitation its partners can observe, and the catalogue holds what CISA can prove and fix.
A probability forecast needs calibration: of the CVEs given 0.2, about 20 percent should turn out exploited. The diagram below simulates a toy world where the truth is known, so you can see good and bad calibration. It isn't a measurement of EPSS.
Mark the ones that apply to you. Three is enough to finish the exercise.
The maths
A forecaster is calibrated when, among all cases it scored s, a fraction s happened. Call y the outcome (1 or 0):
calibration: E[y | score = s] = s, Brier score B = (1 / n) Σ (si − yi)2
P(exploited | flagged) = P(flagged | exploited) P(exploited) / [ P(flagged | exploited) P(exploited) + P(flagged | not) (1 − P(exploited)) ]
Why the points scatter. In a bin of n forecasts with observed fraction p, the standard error is √(p(1 − p) / n), so a calibrated forecaster's points sit near the diagonal, within about two standard errors. An over-confident forecaster says 0.9 and sees far less. An under-confident one says 0.9 and sees more. Why base rates dominate. When P(exploited) is small, unexploited CVEs vastly outnumber exploited ones, so even a small false-positive rate produces many false alarms. Counting the same thing out of 10,000 CVEs makes the arithmetic easy to see.
In practiceAsk any scoring system how it was calibrated, against which outcomes and for which population. A forecast that has never been checked against outcomes is an opinion, and everyone's got one lol.
DefenceCombine evidence, forecast and context, and prefer probabilities you can add up to labels you cannot. When the base rate is low, most of a flagged set is false alarms, however good the ranking.
Where this comes from
- EPSS scores from FIRST, https://www.first.org/epss. EPSS is maintained by the EPSS SIG at FIRST, and the scores are generated by Empirical Security and published freely. Cite: Jacobs, Romanosky, Edwards, Roytman and Adjerid (2021), "Exploit Prediction Scoring System", Digital Threats: Research and Practice 2(3).
- CISA's Known Exploited Vulnerabilities catalogue, distributed under CC0 1.0.
- This page is not affiliated with or endorsed by FIRST, Empirical Security, CISA or DHS.
- The snapshot keeps cumulative counts at a fixed grid of thresholds and every catalogue CVE with its score, percentile, vendor, product and deadline. It also keeps the highest scores outside the catalogue and a seeded sample of the rest.
- Every field was checked when the snapshot was built and again when this page loaded. Text from the datasets reaches the page as plain text, and a CVE becomes a link to NVD only when it has the CVE-YYYY-NNNN form.
- The catalogue's deadlines come from Binding Operational Directives, and BOD 22-01 created the catalogue.
- The estate, its severity labels and the simulated forecasters are synthetic. No real organisation, host or person appears.
- Nothing is fetched: the page runs from the snapshot in this file, and the links here open only when you follow them. The snapshot is rebuilt by
tools/make-epss.js, which refuses any input it cannot validate.