We caught our own AI security scanner making up vulnerabilities
SecFoo is an open-source CLI that runs AI coding agents (Claude, Copilot, Codex, or a plain API key) as security reviewers against your codebase, tracks findings over time, and serves a local dashboard. One of the agent
SecFoo is an open-source CLI that runs AI coding agents (Claude, Copilot, Codex, or a plain API key) as security reviewers against your codebase, tracks findings over time, and serves a local dashboard. One of the agent options — call it secfoo — is supposed to work with nothing but an OpenAI/Anthropic/Gemini API key, no coding-agent CLI installed.
While testing it end-to-end, we found it was fabricating entire security reports.
What we found
We ran it against a small, intentionally-vulnerable Flask app (the whole codebase is one file, app.py) and asked it to do a SAST scan. It came back with "Confirmed" findings — SQL injection, XSS — at files like app/db.py and app/upload.py.
Neither file exists anywhere in the repository.
We ran it again with a different scan type (secret scanning) against the same target. This time it reported a "Confirmed, looks live" AWS access key in deploy/ci.yml, with a recommendation to rotate it immediately.
There is no deploy/ directory in that repository. There is no AWS key anywhere in it.
Why it happened
The root cause was almost embarrassingly simple once we traced it: that agent's code builds the message it sends to the model from the prompt text alone — it never reads the target's actual source files and attaches them. The model is asked "review this codebase for SQL injection, XSS, etc." with no codebase attached. It has no way to say "I don't have the file," because nothing tells it a file was supposed to be there. So it does what a language model does when asked a confident question with no grounding: it generates a plausible, well-formatted answer that matches the shape of a real report.
We'd already built a second, correct implementation (api) that does read and attach the actual files before calling the model — the bug was that the old, broken draft never got deleted when the new one shipped. Both were selectable, with nothing distinguishing a working scan from a hallucinated one. The existing unit test for the broken version actually asserted the broken behavior as correct, so CI never caught it.
What we did about it
- Reproduced it twice, on two different scan types, against the same real target, to rule out a fluke
- Root-caused it to the exact missing step (the broken agent never attaches source files, confirmed via the code and via a passing-but-wrong unit test)
- Deleted the broken implementation entirely rather than patch around it
- Ran the full test suite (700+ tests) plus security linting on the fix
- Opened it as a PR with the full repro steps attached, so anyone can verify it independently: https://github.com/secfoo-com/secfoo/pull/20
Why this matters beyond our own bug
This is the failure mode you should worry about with any "AI reviews your code" tool: a run that reports success, produces a professional-looking report, uses confident language like "Confirmed" and "high confidence" — and is completely disconnected from the actual code. Not a crash, not an error message. A clean green checkmark hiding zero actual review.
We think the fix for that isn't "trust the AI less," it's "verify the AI the same way you'd verify any other system" — reproduce, trace to root cause, and write a test that would have actually caught it. That's what we did here, and it's the standard we're trying to hold every skill and every agent in SecFoo to.
Try it
pip install secfoo
secfoo run --skill sast --agent claude --target https://github.com/your/repo
Repo: https://github.com/secfoo-com/secfoo — MIT licensed, 9 skills (SAST, secret scanning, threat modeling, architecture review, and more), works with Claude, Copilot, Codex, or a plain API key.
We're a small team building this in the open. Issues, PRs, and skepticism all welcome.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.