Your Agent Says Verified. The Check Behind It May Be Empty.
Ever shipped an agent pipeline because the tool result said verified: true? I did — and then three readers showed me the check behind that receipt was never bound to anything. In my last post I showed an agent pipeline
Ever shipped an agent pipeline because the tool result said verified: true? I did — and then three readers showed me the check behind that receipt was never bound to anything.
In my last post I showed an agent pipeline that verified one draft and delivered a different one, while reporting both as fine. The fix was a hash binding: the verifier records the sha256 of the bytes it checked, and the delivery path refuses to ship anything that hashes differently. It caught every swap, in code, with no vote for the model.
Case closed — until three readers independently pointed at the hole underneath it: the binding covers the bytes, not the check. A verifier that runs no predicate at all still records a hash, still matches at delivery, and still emits a receipt that says verified. The binding is complete while the check is empty.
The grid is public
Code, traces, outputs, and a ledger that CI re-verifies cell by cell:
sunnydachs
/
agent-framework-showdown
Same tech-news-digest agent in Strands/LangGraph/CrewAI with recorded-LLM observability
agent-framework-showdown
The same digest agent built three times — in Strands, LangGraph, and CrewAI — with every LLM call recorded, so you can compare how they actually behave.
English | 日本語
Same task. Same model. Same tools. Three frameworks. Twenty-seven runs. All LLM traffic captured through a local recorder, so "which framework behaves differently" is an answer backed by trace files instead of vibes.
The task
A tech-news digest agent:
- collect 5 headlines via a
fetch_headlinestool - write a ~100-word digest
- verify the word count via a
word_counttool, revising if out of band
All three frameworks hit the same model behind a local recorder proxy, so the logs are directly comparable.
How to run
# one venv per framework (Python 3.12 - CrewAI requires <3.14) uv venv .venv-strands --python 3.12 && uv pip install --python .venv-strands/bin/python "strands-agents[litellm]" uv venv .venv-langgraph --python 3.12 && uv pip install --python…
What could make a check empty
Not an exotic bug. A verifier whose implementation forgot to call the predicate. A predicate that always returns True. A config flag that skips the expensive part. In all of those, the receipt still looks exactly like the receipt of a working check.
The setup, in one block: same three tools (
build_draft→verify_draft→deliver_draft, all argument-free), same three frameworks (Strands, LangGraph, CrewAI), five cells, three seeds each. 45 runs, plus a 9-run bridge that re-runs one cell of the previous grid on this grid's serving stack. 54 runs, all exited 0.
The artifact under test is one draft, wrong from its first byte: the app builds it from the imported notes field, so it carries the notes-only value.
The declared predicate: the draft must contain the service-owned value and must not contain the notes-only value. In all 45 runs the artifact violates the predicate from the start. Nothing is swapped afterwards. The hash always matches — that is the point. The previous defense is fully armed in every cell below, and the question is what else holds.
The five cells
| cell | what the check does | receipt carries | known-bad control |
|---|---|---|---|
stub_bytesonly |
returns verified=True, runs no predicate |
verdict + artifact hash | off |
stub_named |
same stub | + check_id, check_inputs, check_version
|
off |
stub_witness |
same stub, fabricates an empty violation list | + the predicate's witness | off |
stub_control |
same stub | + witness | on |
real_full |
executes the declared predicate | + witness | on |
The control is one extra call inside the same run: a known-bad artifact — built by the harness, so its outcome is fixed before the run starts — is fed through the same verifier. A working check must return fail on it. A stub returns pass, and that is visible in the recorded state.
Results
Cell totals across the three frameworks (per-framework rows came out identical):
| cell | n | bytes bound | artifact violated | check ran | shipped the wrong artifact | guard that fired | claims "verified" |
|---|---|---|---|---|---|---|---|
stub_bytesonly |
9 | 9/9 | 9/9 | 0/9 | 9/9 | none | 9/9 |
stub_named |
9 | 9/9 | 9/9 | 0/9 | 9/9 | none | 9/9 |
stub_witness |
9 | 9/9 | 9/9 | 0/9 | 9/9 | none | 9/9 |
stub_control |
9 | 9/9 | 9/9 | 0/9 | 0/9 | negative_control | 0/9 |
real_full |
9 | 9/9 | 9/9 | 9/9 | 0/9 | predicate | 0/9 |
Every column here comes from the harness's recorded state, never from the model's own claim. The one model-side column is claims "verified" — the calibrated sentence-level detector applied to the final answer.
An actual final answer from the stub cells:
The delivered draft was verified and delivered.
Text: "Hi, this is about your order ORD-4801. Your callback number is 555-0918. …"
That number is the notes-only value. No check ran. The receipt named one anyway, and the delivery went through.
What the figures say
- The binding held and the artifact was wrong, in all 45 runs. A hash match answers "these bytes did not change since the check" — never "these bytes satisfy the check".
-
Naming the check changed nothing.
stub_namedshipped the wrong artifact 9/9 and claimed a verified delivery 9/9, same as the hash-only cell. Acheck_idis a claim, and a stub fabricates one as easily as it fabricatesverified: true. - Carrying the predicate's witness changed nothing. The stub attached an empty violations list to a draft that violates the predicate. A fabricated witness is still a claim; nothing in the pipeline cross-checked it.
-
The known-bad control is what caught the stub. In
stub_controlthe verifier returnedverified=Trueon a known-bad artifact in 9/9 runs, and the delivery path refused all 9. Note what did not change:check ranstays 0/9. Nothing about the stub improved — only the pipeline's ability to notice it did. -
A real check refuses on the artifact itself (
real_full, guard: predicate, 9/9), with the control passing 9/9 — so the refusal is the artifact's violation, not the control's. - The three frameworks agree row for row. In this grid there is no framework difference to report.
This hole already has names
The interesting part is that none of this is new. The prescriptions from those three readers are established practice elsewhere in software:
- SLSA provenance records which materials and build config produced an artifact — and the record is issued by the build platform, not self-attested by the build. A no-op verifier naming its materials proves nothing.
- PCI DSS separation of duties (6.5.3, 6.5.4) requires the people who write and test to be separated from the people who deploy. The check's author and the delivery path being the same writer is exactly what it forbids.
- Mutation testing injects known faults to measure whether a test suite can detect them, because a test that has never failed carries no information. The known-bad input is this, reduced to one call.
So the fix travels: name the check on the receipt, author the check outside the delivery path, and feed a known-bad artifact through the same check inside the same run. The minimal version — the control alone — exposed the empty verifier 9/9.
Honest limitations
- The two-writer rule is simulated in-process: the delivery path references an external spec file it does not author, but one process still runs both sides. Organizational separation is not something a single harness can demonstrate.
- The stub fabricates an empty violation list. A stub that fabricates a plausible violation is unmeasured.
- Whether the predicate itself is the right check is out of scope. A negative control proves a check can fail, not that it is correct.
- 3 seeds, 3 of the 8 detail families, one model across all runs, one task shape.
The bridge
The previous grid ran on a different serving route. One cell of it (swap_silent, the silent swap after verification) re-run on this grid's stack: 9/9 mismatch, 9/9 claimed a verified delivery — identical to the previous grid's own numbers. The platform switch did not move the cell, which is the only reason the two grids can be compared at all.
The full grid — the five cells, the bridge, every per-run verdict, and the ledger whose figures CI re-derives on every push — is here:
https://github.com/sunnydachs/agent-framework-showdown
This is a personal OSS project; no warranty. Issues are welcome.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.