Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Let the model decide — then read the audit

Most "governed agent" demos are governed because the demo author pulled every string. The agent calls the tool the script says, in the order the script says, and the audit comes out clean because nothing could have gone

Most "governed agent" demos are governed because the demo author pulled every string. The agent calls the tool the script says, in the order the script says, and the audit comes out clean because nothing could have gone another way. That proves the plumbing. It does not prove governance, because governance is what happens when the thing you are governing has a will of its own.

So we took the strings off.

Agenthof is a small control plane for AI agents with one rule: an agent gets capability only through doors — model, tool, exec, spawn — and each door authorizes the call, injects the credential the agent never holds, and writes a ledger event bound to the human who started the run. The reference agent (LangChain, Python) used to drive those doors from a fixed script. Now a real model is handed the agent's actual capabilities as tools — the exec door, the spawn door, and exactly the tools its grant mirrors — and decides, round by round, what to call.

No new door was needed. The model door authorizes only the logical model name, rewrites it to the provider's, and forwards the rest of the tool-calling conversation as it is. So an agentic loop is just more governed model calls, plus whichever exec, tool and spawn calls the model chose, each a first-hand event of its own. The agent is untrusted; the doors govern it regardless.

What the audit shows

Run it against your own OpenAI-compatible provider. The script prints agenthof audit <run> for the parent run. Its header says ledger integrity: verified (N events); below it, in order: a model fast — N prompt / M completion tokens line per round the model took; an exec event the operator's runtime attested first-hand; a tool call made as the human through a token exchange, and another through a bridge runtime that attests it ran the tool; and two spawn events, each a child run with its own ledger, linked to the parent and bound to the same person.

If the model asked for something it was not granted, there is a recorded refusal. When the exec or spawn door refuses, the agent tells the model refused by policy; when the tool door refuses, the model is told the door call failed. Either way the model decides what to do next, and the step may still succeed. A refused model call instead fails the step — there is no model to ask. Against a real provider that is the likeliest refusal: a rate limit is recorded as model fast refused — budget.

agenthof investigate --run <run> then prints the whole tree as one timeline, each child's events indented under the parent's. The script runs audit verify on the parent and each child, and removes its working directory on exit. Your key is removed from the environment before any process starts and reaches the host agenthof process per command only; the agent never sees it.

What it does not show

Three honest limits.

It does not show containment. The compartments here are host processes started by test stand-ins. Whether an agent could reach a provider on its own is the operator's runtime's job; the project keeps separate, podman-based proofs for that.

It does not promise an outcome. Real models vary; a run may skip a door, call one twice, or take another path. The promise is narrower: whatever the model did was governed and is on the ledger. The deterministic proof of the loop is a hermetic CI run that plays a fixed tool-calling script against the same doors.

And the agent's own report is not evidence. The step artifact — exec: …, tool whoami: … — is the untrusted agent summarizing itself. The evidence is the door events, written by Agenthof and the runtimes. If they ever disagreed, the ledger is what happened. The ledger is append-only and hash-chained: it catches an edit, deletion or reordering that does not recompute every later hash. It is tamper-evident within those limits, not tamper-proof, and the docs say exactly where the limits are.

Any framework

The agent is two layers: a framework-neutral door layer — plain HTTP and MCP against a published wire contract, which also defines the two door tools and discovers the granted tools — and a thin LangChain shell that binds those tools and carries each model turn. Porting to another framework means rewriting the shell. The doors do not know or care which framework is talking. First-class support for agents built on other runtimes and frameworks is on the roadmap as "More harness adapters"; the point of this one is that it needed no special treatment.

The showcase, the hermetic CI run, and the honest-limits page are in the repo under Apache-2.0 (tagged v0.1.0). Bring a key, run it, and read the audit — then decide for yourself what it proves.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.