Your detector's threshold is a benign-only quantity
A guardrail's threshold looks like a model parameter. It isn't. It's a property of your traffic — and there's a one-line proof, which matters because the thing most people calibrate it on is the wrong dataset. I measure
A guardrail's threshold looks like a model parameter. It isn't. It's a property of your traffic — and there's a one-line proof, which matters because the thing most people calibrate it on is the wrong dataset.
I measured this on a public benchmark of 629 real prompt-injection attacks (AgentDojo payloads buried inside ordinary tool output — bills, emails, web pages) plus 97 benign tool outputs, over nine open-source detectors. The benchmark ships the raw score for every detector on every sample, so this is a re-measurement of published data, not a new experiment.
The default 0.5 fails in both directions
Prompt Guard 2's scores on that data: attacks around 0.009, benign around 0.0008. The decision cutoff is 0.5 — roughly 50x above the model's entire range. It catches 6 of 629 attacks (1.0%) and never fires on benign traffic.
That's the famous failure: a detector that is running, returning valid scores on every request, and configured to catch nothing.
The opposite failure is in the same table. Two detectors in the set score benign traffic at ~0.999. For them a 0.5 cutoff sits below their entire range:
| detector | TPR @ 0.5 | FPR @ 0.5 | benign median |
|---|---|---|---|
| prompt-guard-2-86m | 1.0% | 0.0% | 0.00075 |
| protectai-deberta-v2 | 23.1% | 4.1% | 0.00003 |
| jailbreak-detector-large | 50.7% | 2.1% | 0.0038 |
| testsavant-defender | 58.8% | 48.5% | 0.43 |
| preamble-defense | 88.4% | 47.4% | 0.24 |
| deepset-deberta | 100% | 97.9% | 0.999 |
| fmops-distilbert | 100% | 97.9% | 1.0 |
One constant produces "catches nothing" and "screams at toast" inside the same benchmark, because the score scales span five orders of magnitude. A default threshold is not a default — it is an assumption about a scale your model may not share.
The threshold is a benign-only quantity
Suppose you want the false-alarm rate at or below some budget f. The achievable operating points are set by the benign score distribution alone: the maximum-TPR threshold is the k-th highest benign score, with k = floor(f · n_benign). The false-alarm rate is a function of the benign scores only, and TPR is monotone in the threshold — so no labelled attack can move the optimal threshold for a given budget.
The attack labels choose the budget. Benign traffic sets the threshold.
I checked that against the data rather than asserting it. Sweeping an attack-labelled threshold — maximising TPR subject to FPR at or below each detector's own benign-derived FPR — reproduces the benign-derived TPR exactly, to the decimal, for all nine detectors: 98.7 / 33.2 / 14.1 / 47.5 / 48.8 / 6.2 / 0.0 …
So "calibrate the threshold against real attack traffic" is a category error. Attacks are how you measure the payoff. They are not the knob. (You can absolutely use labelled attacks to decide which budget is worth paying for — that's a cost decision. It still doesn't change where the threshold sits.)
The part that bites: it doesn't transfer
If the threshold is a benign-only quantity, it follows that it is only correct for the benign traffic you measured. That is the trap.
Take the same benchmark, split into four domain suites, and calibrate the threshold on three of them at a 2% false-alarm budget. Apply that threshold to the held-out fourth. It breaks the budget on 11 of 36 folds — the held-out false-alarm rate is 4.9%, 2.5x what was promised — and the offenders are exactly the folds whose benign scores sit an order of magnitude higher.
- prompt-guard-2-22m: benign median on travel is 0.0092 versus ~0.0025 elsewhere. The threshold carried in from the other folds flags 13 of 20 benign samples there (65%).
- prompt-guard-2-86m on slack: 5 of 21 flagged (24%), because slack's highest benign score is 4x travel's.
The detector didn't change. The traffic did. Because the threshold is derived from benign traffic, it moved with it — and the calibration you did last quarter is now wrong by a factor of five.
What to do with this
- Derive the threshold per traffic source, or per rolling window — not per model. Ship a calibrator, not a constant. The number belongs to the deployment, not to the checkpoint.
- Monitor the benign score distribution. Its median and p99 are the leading indicator: when they drift, your effective false-alarm rate has moved even though nobody touched the config. The attack side only tells you the payoff, after the fact.
- Re-derive at an n that supports it. A 2% budget needs a few hundred benign samples before the quantile means anything; below that, the "budget" is a rounding artefact.
- Keep the failure visible. This failure is dangerous because it has no symptom: the guardrail is up, healthy, returning valid scores. If nothing reports the rate, "green" and "catching nothing" look identical from the dashboard — which is equally true of an eval that only inspects a quarter of the system it claims to grade.
None of this needs labelled attack traffic in production. It needs you to know what your normal looks like — and to re-check it when normal changes.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.