Prove the Guardrails Worked: Turning a LangChain Agent's Run Into a Replay Bundle
A CISO asking about an AI agent incident does not ask what the agent was supposed to do. They ask what it actually did, on this run, at this time, against this policy. Answering that by grepping through every server's ap
A CISO asking about an AI agent incident does not ask what the agent was supposed to do. They ask what it actually did, on this run, at this time, against this policy. Answering that by grepping through every server's application logs is not an answer. When the bosses want answers, it's not time for a research project.
Cognous's Open Control Stack exists to make that question answerable in minutes instead of days. It organizes agent governance into four layers: declare what an agent may do, control what it actually does at runtime, replay any specific run after the fact, and package the result as evidence a reviewer can act on.
Wiring the Agent Action Manifest into LangChain covered the first layer: a manifest that declares which actions a LangChain agent may take, with anything undeclared falling to a default posture of escalate. Wiring LangChain into the Agent Control Plane covered the second: routing every one of those decisions through RunRecorder as wrap_tool_call middleware around a real create_agent loop, so the run accumulates a real record as it happens.
That record answers "what did the policy decide." It does not yet answer "can I hand this to an auditor." A RunRecorder export lives inside the Control Plane's own object model. Turning it into something portable, independently verifiable, and safe to hand outside the team that built the agent is a different problem: the Replay layer.
The Run So Far
The data pipeline agent from the last two posts maintains schema changes against an internal analytics database, using the same manifest and the same manifest_guard middleware built in the Control Plane post. Its manifest declares two actions: schema_add_column, which needs no review, and table_drop, which needs human approval because dropping a table is destructive and irreversible. A third call, renaming a table, was never declared at all.
The agent runs through a real create_agent loop, with the guard intercepting every tool call before it executes:
agent = create_agent(
model=fake_model,
tools=[schema_add_column, table_drop, schema_rename_table],
middleware=[manifest_guard],
)
result = agent.invoke({"messages": [HumanMessage("Apply the pending schema changes.")]})
Running it against the real agent-control-plane and langchain packages produces:
tool result -> added column loyalty_tier (text) to customers
tool result -> block: Tool 'table_drop' is explicitly blocked in this frame.
tool result -> escalate: Tool 'schema_rename_table' is not in the allowed-tools list and requires manual review.
agent -> Finished the requested schema maintenance.
schema_add_column allows because it's declared, needs no review, and the run carries the write authority it requires. table_drop blocks outright: it's declared, but its approval_required review mode puts it in the Control Plane's own blocked list, so the gate stops it before it runs rather than letting it fall through. schema_rename_table escalates for a different reason: it isn't in the manifest at all, so it was never in the frame's allowed list, and the gate's own default catches it without anyone having to name "renaming tables" as a risk in advance.
Calling recorder.generate_replay_bundle() produces a ReplayBundle object. A ReplayBundle might sound like a finished artifact, but it is not.
LangChain's part ends here. Once agent.invoke() returns, there's no more agent loop, no more tool calls, no more messages, just a plain JSON export sitting on disk. Everything from this point forward treats that export as a record to validate, redact, sign, and verify, the same as it would for a run produced by any other framework.
The Bundle That Isn't a Bundle Yet
Cognous's Replay layer, arb, defines its own schema for a portable replay bundle: AgentReplayBundle. It exists independently from the Control Plane. That independence is the point of a replay layer: it can validate, redact, sign, and verify a run record from any bundle, regardless of its origin. But this leads to a complication: the Control Plane's own export doesn't line up with arb's schema. The fields are mismatched:
-
Identifier:
replay_bundle_idversusbundle_id -
Proposals:
actionsversusaction_proposals -
Decisions:
decisionsversuspolicy_decisions
This is not a bug in either repository. The Control Plane's export is an internal snapshot of its own recorder state. arb's schema is a public interchange format meant to outlive any one producer. A mapping step between them is expected, not a workaround:
arb_bundle = {
"bundle_id": cp["replay_bundle_id"],
"bundle_version": "0.1",
"run_id": cp["run_id"],
"status": "complete",
"generated_at": cp["generated_at"],
"producer": "agent-control-plane",
"frame": cp["frame"],
"action_proposals": cp["actions"],
"policy_decisions": cp["decisions"],
"policy_traces": cp["policy_traces"],
"blocked_actions": cp["blocked_actions"],
"authority_records": cp["authority_records"],
"reliance_records": cp["reliance_records"],
"final_output": cp["final_output"],
}
With the mapping applied, the file validates cleanly:
$ arb validate pipeline_agent_bundle.json
VALID bundle_id=138eb7ca-3da9-4de8-b197-57062d8fcf00 issues=1
WARNING W012: signature_metadata is missing.
One warning left: the bundle isn't signed yet.
From Valid to Handoff-Ready
A valid bundle still contains raw payloads, targets, and reasons from an internal database run. Before this goes to anyone outside the data platform team, redaction strips those out:
$ arb redact pipeline_agent_bundle.json --out pipeline_agent_bundle.redacted.json --targets --reasons --final-output
Redacted bundle written to pipeline_agent_bundle.redacted.json
Every action's payload is redacted unconditionally, on every run of the command. --targets, --reasons, and --final-output are optional additions:
-
--targetsredacts every action'starget -
--reasonsredacts thereasonfield on every action proposal, policy decision, policy trace rule, and blocked action -
--final-outputredacts the run'sfinal_output
What survives untouched: action IDs, timestamps, tool names, and the decision results themselves. A reviewer can still see that table_drop was blocked at a specific time under a specific policy version. They just can't see the table name.
The bundle's status field flips from complete to redacted, and it still validates:
$ arb validate pipeline_agent_bundle.redacted.json
VALID bundle_id=138eb7ca-3da9-4de8-b197-57062d8fcf00 issues=1
WARNING W012: signature_metadata is missing.
Signing comes after redaction, not before: it proves the exact bundle a reviewer opens is the one the data platform team produced, unmodified since.
$ arb sign pipeline_agent_bundle.redacted.json --secret pipeline-signing-secret --key-id data-platform-signing-key --out pipeline_agent_bundle.signed.json
Signed bundle written to pipeline_agent_bundle.signed.json
$ arb verify pipeline_agent_bundle.signed.json --secret pipeline-signing-secret
Signature VALID
Once signed, the file can not be validated. arb verify is the only command built to read a signed bundle. Treat the signed file as the thing that gets handed off, not the thing that gets edited.
What This Actually Proves
This agent proposed three actions. The policy evaluated each one under this exact frame and authority set. table_drop never had a chance to run because its review requirement put it in the blocked list before the agent's call was even evaluated. The signature proves this exact bundle hasn't been altered since the data platform team produced it.
That's the point of the Replay layer: every agent workflow produces a signed, trusted artifact of what it did, not just what it was supposed to do. During an outage, a root cause investigation, or a postmortem, that record is the difference between reconstructing an agent's behavior from memory and fragments of application logs, and pulling up exactly what happened, signed and ready to hand to whoever's asking.
Clone Agent Replay Bundle and try arb against the example bundles in the repo, or your own LangChain agent's exports. Learn more about Agent Replay Bundle at cogno.us.
The next piece takes that bundle and turns it into something a business reviewer can read without touching JSON: the Agent Governance Evidence Pack.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes â full credit and traffic to the original publisher.