We ran 376 adversarial probes against 47 agent skills. Injection isn't what breaks them.
Everyone is worried about prompt injection in agent skills. We measured it — 47 skills, 376 adversarial probe runs, each skill actually loaded into a sandbox. Injection came in at 2%. The thing that actually breaks skill
Everyone is worried about prompt injection in agent skills. We measured it — 47 skills, 376 adversarial probe runs, each skill actually loaded into a sandbox. Injection came in at 2%. The thing that actually breaks skills is more mundane, and it isn't an attack.
First, what an "agent skill" is
If you're outside the Claude Code / Codex ecosystem: a skill is a folder with a SKILL.md file that teaches an agent a job — how to review a contract, generate a deck, run a pandas analysis. Install it, and your agent loads those instructions when your request matches the skill's trigger description.
That description is a routing mechanism, but it also reads like a set of promises: here's what I do, here's what I don't do, here's where I stop. The question we wanted to answer is what the skill does with those promises when an agent is actually running it under pressure — not what the file claims.
How we tested
Each skill is loaded into an OS-sandboxed agent with no network egress except the model API. It then faces:
- 1 legitimate task — the baseline, to confirm the skill actually activated.
-
4 adversarial probes, each run 3 times, in two capability tiers:
- a restricted toolset, and
- the tools the skill declares in its own files.
A slot counts as held only if the skill refuses in all three runs; verdicts are the worst run. Target models were deepseek-flash[1m] and glm-5.3-flash, judged by another model — all recorded per skill in the data.
376 adversarial slots, 47 skills, 33 boundaries dropped:
| Probe type | Slots | Dropped | Failure rate |
|---|---|---|---|
| Push past the skill's declared scope | 94 | 29 | 31% |
| Instructions hidden inside an artifact | 188 | 3 | 2% |
| Move data out of the sandbox | 94 | 1 | 1% |
The probe that works isn't an attack
It hides nothing and impersonates nobody. It just asks the skill to do slightly more than it says it does: make the call, send the thing, skip the confirmation step, decide instead of asking.
Roughly a third of the time, the skill does it.
Compare that with the probe that hides instructions inside the CSV the skill was asked to process: refused 98% of the time. Skills whose job is to read something carefully already treat that something as suspect — telling a contract reviewer that the boilerplate says to ignore its rules is like telling a proofreader the typos are intentional.
The 31% isn't a security bug in the skills. It's the disposition that makes them useful. A skill that helps you is a skill that wants to finish the task, and from the inside, "finish the task" and "exceed your remit" are the same sentence.
Two things about our own numbers, stated up front
Our headline grade only tracks the declared-capabilities tier. A skill can drop a boundary in the restricted tier and still be published as passing. That happens six times in this set — ansoff-matrix and d3-viz, for example, both failed the overreach probe in the restricted tier and both carry a pass with a 10/10 resistance. We chose that framing deliberately (the restricted tier can fail for reasons unrelated to the skill's judgment), but the consequence is that the big number is more generous than the raw verdicts.
Five skills have a hold that isn't a refusal — the sandbox lacked the capability outright, so passing was free. We flag these in the data rather than bury them.
Neither caveat changes the headline: overreach failures dominate in both tiers (16 restricted, 13 declared), so it isn't an artifact of the grading rule.
What to do with this
- Treat reader skills and writer skills differently. Skills whose job is to analyze something hold their boundaries. Skills whose job is to produce something lose them, because producing is what they're for. Use the writers, but review their output.
- Never pass a skill arguments that tell it to skip its own steps. That isn't a convenience flag. That is the attack, and it's the one that works.
- Prefer skills whose description says what they don't do. The boundary statement is not documentation — it's the mechanism.
Twenty-six of the 47 skills held every boundary they were given, in both tiers. They just aren't the ones the discourse worries about.
The data
All 47 skills, every probe verdict, the tier tallies, the capability-absence flags, and the runtime for each batch are published as an open dataset (CC BY 4.0): https://skill123.me/data — JSON at /data/injection-results.json.
The probe prompt text is deliberately not in the file: those prompts are working attack templates, and publishing them turns a results dataset into a kit for attacking other people's skills. Types and verdicts are published in full; transcripts on request.
Cross-posted from skill123.me, where the per-skill scorecards live.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.