There are three ways a security guardrail can fail you.
Two of them are loud. One of them is a serial killer.
- It's down. Crashed, misconfigured, not deployed. You find out fast — errors, alerts, a red dashboard, an on-call page at 3am. Painful, but honest.
- It's weak. It runs, it catches some attacks, it misses others. You can measure it, argue about it, improve it. Also honest.
- It's running perfectly, passing every health check, returning valid scores on every request — and configured to catch nothing. No error. No alert. No red anything. Every dashboard it touches is green by construction.
The third one is the problem. Because it has no symptom. Your monitoring says "guardrail: healthy ✅" and it's telling the truth. The guardrail is healthy. It's just not a guardrail.
I ran into a textbook case of #3 while benchmarking prompt-injection detectors, and it's worth dissecting, because once you see the shape of it you'll start finding it everywhere.
The setup: the smoke alarm of the AI stack
Everyone shipping an AI agent bolts on a prompt-injection detector — a little classifier that reads the text flowing into the model and screams if it smells an attack. It's the smoke alarm of the agent world. You install it, you see the green light, you move on.
So I did the obvious experiment: I took 10 of these smoke alarms and set 629 real fires.
Specifically — I took 629 real injection attacks from AgentDojo, buried each one inside ordinary tool output (a bill, an email, a web page — the way an agent firewall actually sees them, not the clean lab version), and ran 10 open-source detectors over the lot, plus 97 benign tool outputs to catch the ones that just alarm at everything.
Most of the results were the boring kind of bad (the "weak" and "screams at toast" failures). But one detector produced the interesting kind of bad.
Exhibit A: Meta's Prompt Guard 2, catching 1%
Meta's Prompt Guard 2 — the model everyone name-drops — caught 6 of 629 buried attacks. About 1%.
Now, your first instinct is "the model is bad." Reasonable instinct. Wrong.
Here's what the model was actually doing on each request:
| Attack | Benign | |
|---|---|---|
| Prompt Guard 2 score | ~0.009 | ~0.0008 |
| The 0.5 decision cutoff | 0.5 | 0.5 |
| Verdict | ✅ ALLOW | ✅ ALLOW |
Look at the top row. The attack scores roughly 10x higher than normal text — the model's ranking is nearly perfect. As a discriminator, it works.
Now look at the cutoff. Use the model the obvious way — the standard 0.5 decision boundary you'd apply to any binary classifier out of the box (block if P(malicious) ≥ 0.5) — and every score it produces, the attack at 0.009 and grandma at 0.0008 alike, sits far below the line. So the verdict on every single request is identical: "nah, we're good." ✅
(To be precise, because this is the sentence people will poke at: 0.5 isn't some evil value Meta hard-coded — it's the threshold you get by treating a probability classifier the normal way. The failure isn't a bad config someone shipped; it's that **the obvious, default way to use this model is off by ~50x for buried attacks* — and nothing about the running system would ever tell you.)*
The detector was producing different scores. The control just wasn't doing anything with the difference.
It's a smoke alarm with the sensitivity dial turned all the way to "only trigger for a literal supernova." The sensor is fine. The wiring is fine. It will faithfully detect the heat death of the universe. Your kitchen fire? Not so much.
That is not a weak guardrail. That is a perfectly functional discriminator thresholded into a no-op. It loads, it returns scores, it passes health checks, it logs clean — and it catches 1% of attacks. Green by construction.
The plot twist that makes it worse
Here's the part that should genuinely unsettle you.
I re-ran Prompt Guard 2 with the threshold moved from 0.5 down to 0.003 — into the range where its scores actually live — and calibrated so it wrongly flags at most 2% of normal traffic, measured on an AgentDojo domain it was never tuned on.
It went from catching 1% to catching 99% (621/629).
Same model. Same weights. Same requests. One number. The difference between "protects you from basically nothing" and "catches almost everything" was a config value the model shipped with — off by roughly two orders of magnitude for this use case.
The uncomfortable takeaway isn't "the model is good" or "the model is bad." It's that the shipped default was wrong by ~100x, and nothing about the running system would ever tell you. Every health check passes at 1% exactly as it does at 99%.
(Honesty clause, because this is the part everyone screenshots wrong: don't read "99%" as "Prompt Guard 2 solves prompt injection." Every AgentDojo attack uses the same wrapper template, so a threshold tuned that finely may be recognizing the template, not attacks in general. The load-bearing claim is the miscalibration — "the default is wrong by 100x" — not the 99%. Tune on **your* traffic, then trust a number.)*
Why "green by construction" is the dangerous class
Compare the three failure modes by how you'd find out:
| Failure | Dashboard | How you learn about it |
|---|---|---|
| Guardrail down | 🔴 red | Instantly — errors, alerts, paging |
| Guardrail weak | 🟡 measurable | Eventually — an incident, a red-team, a metric |
| Guardrail green & useless | 🟢 green | Never — until the breach, and even then you'll swear it was on |
The first two fail toward visibility. The third fails toward silence. It's the security equivalent of a bouncer who's clocked in, in uniform, standing at the door, checking every ID — with his eyes closed. Every metric you have says "door: staffed." The metric you don't have says "staffed ≠ guarding."
And "it passed the health check" is doing a lot of unearned work in most teams' heads. This is the gap between a liveness check (is the thing running?) and an effectiveness check (is it actually stopping what it's supposed to stop?). A health check proves liveness. It says nothing about effectiveness — and almost every dashboard measures the first while quietly implying it measured the second.
How to actually tell if yours is armed
This isn't a "shame on Meta" post — the model does what it was trained to do, and defaults have to be conservative to avoid blocking everyone. It's a "shame on us if we trust the green light" post. Concretely:
- A score that ranks well with a bad cutoff is worth zero. Separately measure discrimination (does it rank attacks above benign?) and operating point (is the threshold where the scores actually live?). A model can be great at the first and useless at the second — which is exactly this case.
- Never trust a vendor's default threshold on your data. It was picked for a distribution that isn't yours. Sweep it. Find where your attack scores and your benign scores actually sit.
- Measure at a false-alarm budget, not in the abstract. "Catches 99%" is meaningless without "…while wrongly blocking X% of normal traffic." A detector that blocks 98% of legit calls isn't a control; it's a way to make everyone route around your agent. (Two detectors in my benchmark did exactly this and still "scored" 100% on attacks. A brick over the deny button scores 100% too.)
- Health check ≠ armed check. Add a test that fires a known attack through the live guardrail and asserts it's blocked. If your monitoring can't tell the difference between "catching 99%" and "catching 1%," your monitoring is green by construction too.
- Red-team the config, not just the model. The vulnerability here wasn't in the weights. It was in a single threshold. Your attack surface includes your YAML.
Or, as a checklist you can staple to any guardrail:
Liveness check: Is the detector running at all?
Effectiveness check: Does it BLOCK a known attack, right now, in prod?
Calibration check: Are attack scores actually separated from benign scores?
Operating-point check: At our tolerable false-alarm rate, what % of attacks do we catch?
Regression check: Is all of the above still true after the next model/config change?
Only the first line is on most dashboards. The other four are the difference between a control and a decoration.
The bigger, more uncomfortable question
Prompt Guard 2 at 1% is a clean, measurable example because I had 629 attacks to throw at it. But the shape generalizes way past prompt injection:
- The WAF rule set that's deployed but in "log-only" mode.
- The alert that's been firing into a muted channel since Q1.
- The MFA that's enforced… except for the legacy API path.
- The test suite that's green because half of it is
skip.
All green. All healthy. All protecting you from a supernova and nothing smaller.
So the question I'd actually sit with: how many of your green dashboards are green because the thing works — and how many are green by construction? The scary answer is that, by definition, you can't tell the difference by looking at the dashboard. You have to fire a real attack at it and watch.
The whole benchmark — 10 detectors, 629 attacks, the threshold sweep, the cross-domain calibration — is open and reproducible here, if you want to see which alarms are real and which are decorative:
github.com/rudratoshs/buried-injections ⭐
The full thing — 10 detectors, 629 attacks, the threshold sweep, the cross-domain calibration — is open and reproducible. If you find a flaw in the methodology, I genuinely want to know — security benchmarks are more useful when people try to break them. And if it made you want to go check whether your own guardrail is armed or just green, a ⭐ helps it reach the next person about to trust a green light. 🙏
Be honest in the comments: have you ever found a security control that was "on," passing every check, and doing absolutely nothing? What was it, and how did you finally catch it — an incident, a red-team, or dumb luck? 👇
I write about AI, security, and the honest ways things break — benchmarks with the false-positive column left in. Follow me here if that's your lane. 👋
Top comments (13)
Ran
bench/at_budget.pyover your own saved scores (bench/results/at_budget_2pct.json) before writing this, because the port is what makes the rest checkable — all nine detectors reproduce both saved columns exactly (in-sample and cross-domain) once the suite rotation is ordered workspace/travel/banking/slack, so the numbers below are yours, not a re-derivation.Three things the pooled figure doesn't say.
1. The 2% budget is an in-sample property; the unseen column is a different quantity.
caught_at_budgetenforces the budget when picking (allowed = floor(budget * len(calib)), per fold on the calibration split), then measures on the held-out suite, where it can exceed it — and does, for eight of the nine (the regex baseline is the exception): 5.2% for Prompt Guard 2 86m, 13.4% for the 22m, 9.3% for testsavant. So "99% at a 2% false-alarm budget" splices an in-sample budget onto a cross-domain TPR. Both are honest numbers; the sentence that gets screenshotted has combined two.2. The pooled TPR hides a fold spread as large as the rate it reports. Per held-out suite, prompt-guard-2-22m is 100% on travel, 22.9% on workspace, 16.0% on banking, 1.0% on slack — pooled 35%, spread 99 points. jailbreak-detector-large runs 17.1% → 100% (pooled 51%). For six of the nine the spread is at least the pooled rate itself.
cross_domainalready has each fold incand accumulates it away (caught += c); keeping the list and printing min/max is two lines, and it changes what "the number to trust" means.3. For a control, the worst fold is the number to trust — not the mean. That's your own thesis one level down: a pooled average is a liveness statistic (it says the detector ran on four suites), while the min fold is closer to an effectiveness one. Prompt Guard 2 86m is the reassuring case — 97–100%, a 3.3-point spread — and most of the field isn't.
The benign-sample size and the rotating-canary points above are the other half of it; both are cheap.
This is exactly the kind of comment I was hoping this post would attract.
You’re right on the 2% point — I compressed two different quantities into one sentence. The budget is enforced during calibration, while the reported held-out/cross-domain catch rate is a different measurement. Putting those side by side without making that distinction explicit makes the headline number easier to screenshot than to interpret.
The fold spread point is even more important. A pooled TPR can look reassuring while one domain is basically a miss. I like the idea of keeping the per-fold values visible and reporting the worst fold alongside the pooled number.
That actually fits the thesis of the article better: the number isn't the measurement unless you preserve the conditions under which it was produced.
I'm going to fix this rather than defend the original wording. Thanks for actually running the code and checking it against the saved results.
This is a really good catch.
I was treating “0 false positives out of 97” as if it established a 2% false-alarm rate, when statistically it really only tells us what happened in that small sample.
Your point about the benign denominator is exactly the kind of thing that gets lost when a benchmark focuses on the attack-success column. At a threshold this close to the benign-score distribution, the uncertainty around the benign sample can matter more than another decimal place on TPR.
And I really like the held-out rotating-canary idea. A fixed canary can gradually become a memorization test instead of a security test.
So the revised measurement should probably report both:
effectiveness on the held-out pool + confidence around the benign false-alarm estimate.
That’s a much more useful production story than simply saying “99% caught.”
Great addition.
On the min-fold question — same re-run, and the answer changed twice while I was measuring it.
Headline pair: yes. But attach n and an interval to the min fold. Min-fold is by construction the smallest numerator, so it is also the widest interval. Your min folds, Wilson 95%, ordered: 97% [94–98], 26% [19–34], 17% [12–24], 1% [0–5], then five of the nine bottoming out at 0/140–240 → [0–3%]. A bare "17%" reads like a measurement; "17% [12–24], 140 attacks" reads like one.
The surprise: the min-fold leaderboard is mostly a tie. Walking the ranked min folds, 6 of the 8 adjacent pairs have overlapping 95% intervals — Prompt Guard 2 86m is cleanly separated at the top, and below third place the ordering is not resolvable at these n (four detectors are all 0% [0–3%]). So print the interval with the rank, or replace the rank with the interval below the top — a rank implies a separation the data doesn't have.
On readability: the per-suite row is four numbers — that is not what makes it unreadable. Put the 4-cell row under the headline and the full table in
results/*.json; what actually costs a reader is three quantities with different denominators (budget / false-alarm rate / TPR) sitting next to each other unlabelled. Label the denominators and you can afford the detail.One more, same re-run, and it cuts the same way: at this corpus size the 2% budget is degenerate.
allowed = floor(0.02 * len(calib)), and calib is 57/77/81/76 benign cases depending on which suite is held out → 1 in all four folds. So cross-domain, "2%" is not a rate at all; it is "exactly one benign case above the line". Reporting the budget as an allowed count until the benign corpus grows (your ~150+ fix) makes that visible from the header — and it is a third reason those two numbers shouldn't share a sentence.Taking all four together, they collapse into one correction I should have made myself: every headline in that benchmark is a percentage printed over a small integer, and a percentage over a small integer implies a resolution the count doesn't carry. That's the post's own thesis one level down — I stripped the conditions (the denominator, the n) that make a number a measurement, and kept the number.
Point by point, because each fix is different:
Min-fold gets an interval and an n, always —
x% [lo–hi], Nper cell. The interval isn't decoration here; it's the thing that stops a min-fold (smallest numerator, widest interval by construction) from being read as a point.The ranking below the top isn't real, so it goes. This is the one that stings and it's the most important. If 6 of 8 adjacent pairs overlap at 95%, the leaderboard is asserting an order the data can't resolve — and a rank is itself "green by construction": it looks like information and encodes none below Prompt Guard 2 86m. So the honest output is a partial order — one detector cleanly separated at the top, and an unranked pack of "indistinguishable at this n" (four literally 0% [0–3%]). Intervals, not positions, below first place.
The unreadability was unit-collision, not count. You're right — four numbers is fine; three different kinds of rate (budget / false-alarm / TPR) sitting unlabelled side by side is the cost. Label the denominators and the detail pays for itself: 4-cell row under the headline, full table in
results/*.json.The 2% budget is an integer wearing a percentage.
floor(0.02 × {57,77,81,76}) = 1in every fold — so "2%" is "exactly one benign over the line," and printing it as a rate invents precision the calibration split doesn't have. It reports as an allowed count until the benign corpus clears the size where the percentage means anything.The question that follows is the one I actually don't know: below the top, is the pack resolvable at any feasible n, or is "these are indistinguishable" the real scientific finding? At the min-fold widths you listed, separating two detectors ~10 points apart needs the attack corpus to grow a lot — and if ranking the pack is just noise-mining, the benchmark's honest output is "one detector clears the bar, the rest are tied," full stop. Did your re-run give you any feel for whether more attacks would ever break the tie, or is the pack genuinely level?
Keeping the false-positive column in, and the honesty clause about the shared AgentDojo template, is what makes the 1% → 99% result believable.
Two things I'd add to the checklist.
The false-alarm budget needs enough benign samples to back it. If "wrongly flags at most 2%" is checked on something like the 97 benign outputs, the data can't really say 2%: with 0 flags in 97 the 95% upper bound on the false-alarm rate is about 3.7%, and with 1 flag it's about 5.6%. To show "at most 2%" with zero flags you need roughly 150 benign samples, and more if a few flags are allowed. At a threshold like 0.003, sitting that close to the benign scores, this is the number most likely to surprise someone in production.
On check 4, the known attack fired through the live guardrail: if it's always the same one, it can keep passing after the detector has drifted, for the same reason you flag with the 99%. It may be recognising that one string. Drawing the canary from a held-out pool of attacks that were never used to set the threshold, and rotating it, turns the armed check into a small ongoing measurement instead of a single fixed probe.
@rudratosh thanks, and a small mix-up: your replies to pm25coder and me seem to have swapped places. The at_budget.py run and the three points about the splice, the fold spread and min-fold are pm25coder's; the benign sample size and the rotating canary were mine. Worth crediting them in the changelog under the right name.
On your canary question: I'd do both, at two speeds. Resample one canary from the held-out pool on every run, as you lean towards, and track the catch rate as a running count with its interval, so drift shows up as a trend. Then on every deploy or threshold change, run the whole held-out pool, which gives a proper effectiveness number at the moment it's most likely to break. And retire any canary that ever gets used to tune anything, or the pool slowly turns back into the fixed probe.
Absolutely — and thanks for the correction on attribution. I’ll make sure the changelog credits the
at_budget.pyanalysis and the benign-sample/canary suggestions separately.I like the two-speed approach.
For normal runs, rotate a canary from the held-out pool and track the catch rate over time. Then, whenever the detector is deployed or the threshold changes, run the entire held-out pool to get the real effectiveness measurement.
The “retire any canary that gets used for tuning” rule is especially important. Otherwise the canary quietly stops being a test and becomes another training signal.
That gives me a much better definition of what the live check should be:
the canary detects drift; the held-out pool measures effectiveness.
That distinction wasn't explicit enough in my original post. Really useful comment.
You report the 1%-to-99% swing on Prompt Guard 2 came down to one threshold number, 0.5 to 0.003, on the same weights. The detail worth underlining is where 0.5 came from: it's the demo default that most classifier wrappers inherit, chosen on a score distribution nobody's real traffic matches. We've seen the same shape in production — scores sitting at 0.009 versus 0.0008 look safely separated until attack phrasing shifts and the gap closes, which is why a one-time sweep doesn't hold. The armed check that survives it is re-running the calibration on fresh traffic, not just the canary.
This is the sharper version of my point, and I'm going to steal the framing: a one-time sweep calibrates to a distribution that's already expiring.
You nailed why 0.5 is so dangerous — it's not a value anyone chose for your traffic, it's the demo default that rides in with the wrapper and never gets questioned because the light stays green.
The bit you added that I under-weighted: the 0.009 vs 0.0008 gap isn't a fixed property of the model, it's a property of the current attack phrasing. Shift the phrasing and the gap closes, and a threshold you swept last quarter is now sitting in the dead zone again — silently, with every health check still passing.
So the real control isn't "sweep once, set threshold, done." It's:
Out of curiosity — in your production case, what cadence did you land on for re-calibration, and was it time-based or triggered by the separation metric closing?
Some comments may only be visible to logged-in visitors. Sign in to view all comments.