Dev.to Security πŸ” Cybersecurity πŸ‘ 0 πŸ“– 5 min read

The Login Endpoint That Handed Out Tokens Nobody Ever Checked

I ran the same security test two ways against the same piece of generated code. One method gave a login endpoint a perfect score. The other gave it a 0.75 and flagged the actual reason: the endpoint hands out a valid JWT

I ran the same security test two ways against the same piece of generated code. One method gave a login endpoint a perfect score. The other gave it a 0.75 and flagged the actual reason: the endpoint hands out a valid JWT on every successful login and nothing downstream ever checks it again.

Same spec. Same model. Same four CWE categories. One real run each. The only variable was when the attacker looked.

The setup

I built GAUNTLEX to run two agents concurrently against a specification: a Builder that implements it, a Breaker that attacks it, at the same time, before any code exists for the attacks to anchor to. The pitch has always been that testing the intent catches things that testing the implementation doesn't. I wanted to find out if that was actually true, isolated from everything else the tool does, so I wrote a harness that runs the comparison directly: the same spec, attacked two ways.

Sequential β€” a Builder writes the code first, then a Breaker attacks whatever came out. This is how almost every AI code-review and security tool works today, including most of what gets bolted onto a CI pipeline.

Concurrent β€” the Breaker attacks the specification itself, in parallel with the Builder, before it has any code to read.

The spec: a Flask /login endpoint β€” bcrypt password check, JWT issued on success, audit logging, rate limiting. The attacks: CWE-89 (SQL injection), CWE-79 (XSS), CWE-287 (improper authentication), CWE-306 (missing authentication on a critical function). Fixed set, both runs, same free-tier model, zero dollars spent. The full harness and raw JSON output are public β€” I'll link them at the end so you can run this yourself and get your own numbers, not mine.

What each one found

The sequential Breaker, reading the finished code, scored the login endpoint 1.0 β€” every attack mitigated. Clean pass. Ship it.

The concurrent Breaker, reasoning from the spec alone, scored it 0.75. Three mitigated, one missed:

CWE-287, concurrent Breaker: "JWT Algorithm Confusion Attack (RS256 to HS256)." Verdict β€” missed: "The code only implements JWT token generation (HS256) but lacks any token verification logic; algorithm confusion attacks target verification, which is absent."

Read that again. The attack it tried wasn't even the real problem β€” algorithm confusion targets a verification step. The actual finding is that the verification step doesn't exist anywhere in the generated code. The endpoint issues tokens. Nothing ever checks one.

Here's what the sequential Breaker did with the same CWE-287 slot, looking at the same code:

CWE-287, sequential Breaker: "Hardcoded default JWT secret enables token forgery." Verdict β€” mitigated: "The JWT_SECRET defaults to a cryptographically random value generated at startup via os.urandom(32).hex(), not a static hardcoded string, preventing token forgery via known default secrets."

That's not a bad finding. It's a real thing to check, and it was correctly ruled out. It's also the generic, textbook guess a Breaker reaches for when the only information it has is what's sitting in front of it, and nothing in front of it hints that an entire verification step is missing. You can't flag the absence of something by staring harder at what's present.

Why this happens

A code-anchored Breaker pattern-matches against what exists. It's good at that β€” give it real source and it reasons precisely about the actual attack surface in front of it. What it structurally can't do is notice the attack surface that should exist and doesn't, because there's nothing on the page pointing at the gap. Missing code doesn't announce itself.

A spec-anchored Breaker starts from a different question: given what this endpoint is supposed to do β€” issue a JWT on login β€” what has to be true somewhere downstream for that to be safe? Verifying the token on a later request is table stakes for any JWT auth flow. The concurrent Breaker reasoned toward that requirement and then checked whether the generated code actually met it. It didn't.

Same model, same prompt style, same four CWEs. The only difference is which artifact the attack gets anchored to, and that difference was the entire gap between a false pass and a real catch.

The speed number, for completeness

Concurrent finished in 69.3 seconds. Sequential took 178.1 β€” the Breaker has to wait for the Builder to finish before it can start. 2.57x, if you're counting. It's a real number and it compounds with every attack and every spec you throw at it, but it's not the point of this post. A tool that's faster and still misses the bug isn't an improvement. The timing is a side effect of running the attacks in parallel instead of in sequence β€” the finding is what running them before the code exists actually catches.

This was one run, on purpose

I'm not publishing a best-of-N here. This is the one run the harness produced, raw JSON included, and bench/concurrent_vs_sequential.py --spec examples/demo_issue.md --runs 5 reproduces the comparison end to end on whatever model you point it at β€” including a fresh speedup number and possibly a different finding, since free-tier models aren't perfectly deterministic. If you get a cleaner result on your own run, I'd genuinely like to know; the mechanism (spec-anchored vs. code-anchored reasoning) is the claim, not this specific JWT finding.

I also pointed GAUNTLEX at its own source code with the same free-tier setup β€” no paid key, the exact model gauntlex setup hands most people by default β€” and published that run too, hash and all, so it can be independently re-verified rather than taken on my word.

git clone https://github.com/sanjoy1234/gauntlex
cd gauntlex && pip install -e . && gauntlex setup
python bench/concurrent_vs_sequential.py --spec examples/demo_issue.md --runs 3

If you run this against your own spec and get an interesting result either way, I'd like to hear about it in the comments.

πŸ“° Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.