Dev.to AI 🤖 Ai 👁 0 📖 4 min read

# Don’t Trust the Agent’s “Done”: Verify the System State

AI agents are increasingly allowed to do real work. They deploy applications. They restart services. They write to databases. They call APIs. They trigger automation chains. And when they finish, they usually return some

AI agents are increasingly allowed to do real work.
They deploy applications.
They restart services.
They write to databases.
They call APIs.
They trigger automation chains.
And when they finish, they usually return something like:

“Done.”
The problem is that a successful tool call is not necessarily proof that the expected real-world state exists.
That distinction becomes increasingly important as AI systems receive more execution authority.

Execution and verification are different jobs

Suppose an agent performs a deployment.
The agent reports:

Deployment successful.
Version 4.7.2 is running.

There are several things that could have happened:

  • the deployment command returned exit code 0;
  • an API accepted the request;
  • the orchestration layer reported success;
  • the agent interpreted the response correctly. But the claim we actually care about may be:
Production is currently serving version 4.7.2.

Those are not the same statement.
The first group describes execution events.
The second describes the resulting system state.
A system should not automatically be allowed to prove the second merely by reporting the first.

Treat the agent’s result as a claim

This is the model I use in SCC Runner:

Claim

Authority
↓
Evidence
↓
Verdict

The agent says:

The service is running.

That becomes a claim.
Now we ask:
What source has authority over that property?
For a systemd service, that might be the service mnager.
For example:

ActiveState=active

For a database operation, the database itself may be the authority.
For a deployment, the runtime environment may be more relevant than the deployment tool that initiated the change.
For an external transaction, the external API or ledger may be the relevant authority.
The important idea is simple:
Evidence should come from a source capable of observing the property being claimed.

Why another LLM is not enough

One common pattern is:

Agent A performs the task
↓
Agent B reviews Agent A

That may improve reasoning quality.
But it does not necessarily provide independent evidence.
If Agent A says:

“The database record exists.”
and Agent B only reviews Agent A’s explanation, Agent B still does not know whether the record actually exists.
Both models may be reasoning about the same self-report.
For claims about external reality, verification needs access to the relevant external authority.

Agent
↓
Claim
Database / Runtime / API / Git / Logs
↓
Evidence

The distinction is not about which model is smarter.
It is about where the evidence comes from.

“False” and “not observable” are different states

There is another important distinction.
Suppose we need to verify:

Service X is running.

But the authority cannot be reached.
That does not prove:

Service X is stopped.
``
It means something else:


text
NOT EVALUABLE
AUTHORITY UNAVAILABLE

Verification systems should preserve this distinction.
Otherwise temporary observability failures become false claims, or worse, unavailable evidence becomes accidental confirmation.
Explicit uncertainty is safer than invented certainty.
## From verification to execution gates
Once a claim can be verified reliably, the result can eventually participate in execution decisions.
Conceptually:


text
Claim
↓
Authority
↓
Evidence
↓
Verdict
↓
PASS / BLOCK

For example:


text
Claim:
Deployment is healthy.
Evidence:
expected version present
health endpoint reachable
required instances ready
Verdict:
CONFIRMED
Decision:
PASS

Or:


text
Claim:
Deployment is healthy.
Evidence:
required state cannot be established
Verdict:
NOT EVALUABLE
Decision:
BLOCK

That second behavior is what makes fail-closed verification interesting.
But I would not start there.
## Start in shadow mode
A new verification system should not receive production authority immediately.
A safer adoption path is:


text
SHADOW
↓
ADVISORY
↓
ENFORCED

### Shadow
Observe claims without influencing execution.
Compare the evidence with the checks humans already perform.
### Advisory
Produce evidence and verdicts.
Humans or existing systems still make the final decision.
### Enforced
Selected verification results become execution conditions.
Only after repeated observation should the verifier gain that authority.
Verification software should not ask to be trusted.
**It should demonstrate that it deserves more authority.**
## The goal is not another monitoring dashboard
Monitoring systems answer questions such as:


text
CPU = 42%
requests = 12,431
error rate = 0.3%

Those metrics can be extremely useful.
But verification asks a different question:
> Does the available evidence actually establish the claim required for this decision?
A log entry saying:


text
deployment completed

may be useful evidence.
But it may not be sufficient evidence for:


text
the new version is healthy in production

The claim determines what evidence is relevant.
## Where this becomes useful
This pattern is particularly useful for:
- AI agents with tool access;
- deployment systems;
- autonomous remediation;
- data pipelines;
- workflow automation;
- infrastructure operations;
- multi-agent systems;
- tool chains where one tool calls another tool.
The more layers between the agent and the final system state, the more dangerous it becomes to equate:


text
tool returned success

with:


text
desired outcome exists

## The simplest useful implementation
You do not need to begin with a large verification framework.
Start with one boring claim.
For example:


text
Claim:
Service nginx is active.
Authority:
systemd
Observed:
ActiveState=active
Verdict:
CONFIRMED

Run that verification next to your existing manual process.
Compare the results.
Repeat.
That is a much stronger foundation than beginning by giving a new system permission to block production
## Build systems that can prove what they did
As AI systems become capable of taking more actions, I think we need to separate two questions:


text
What did the agent say happened?
What can the system actually prove happened?



Those questions are often treated as if they were identical.
They are not.
That is the problem I am working on with SCC Runner.
The full canonical article is here:
https://taiwildlab.com/articles/scc-runner-dont-just-take-your-systems-word-for-it/
SCC Runner is currently available in **Free Early Access**:
https://scc.taiwildlab.com/signup
📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.