Dev.to Security 🔐 Cybersecurity 👁 0 📖 7 min read

AI Agents Need More Than Logs: Why I Built NovaFabric

AI agents are becoming more powerful. They can: modify code call APIs execute shell commands change cloud infrastructure read and write files interact with other agents make decisions But there is a simple question

AI Agents Need More Than Logs: Why I Built NovaFabric

AI agents are becoming more powerful.

They can:

  • modify code
  • call APIs
  • execute shell commands
  • change cloud infrastructure
  • read and write files
  • interact with other agents
  • make decisions

But there is a simple question that is still surprisingly hard to answer:

After an AI agent finishes its work, can we prove what it actually did?

Usually, the answer is: not completely.

That is the problem behind NovaFabric.

Logs are useful, but logs are not evidence

Today we have many good observability tools. They can show things like:

Agent started
↓
LLM called
↓
Tool called
↓
API called
↓
File modified
↓
Agent finished

This is useful for debugging. But imagine that six months later something goes wrong. An auditor, security engineer, researcher, or developer may ask:

  • Which model was used?
  • What prompt did it receive?
  • What did the model return?
  • Which tools did the agent call, and with what arguments?
  • Which files changed?
  • Which network services were contacted?
  • What software versions were installed?
  • Were secrets captured?
  • Was the record changed after the run?
  • Can we reproduce or inspect the execution again?

A traditional trace does not necessarily answer all of these questions. And more importantly, a normal database record can potentially be changed later.

That is where the idea of execution evidence becomes useful.

Think of an AI run as an artifact

NovaFabric takes a different approach. Instead of thinking:

AI Agent → Logs

it thinks:

AI Agent
   ↓
Execution
   ↓
Portable Evidence

NovaFabric captures a run into something called a Run Capsule: a structured folder describing one execution.

capsule/
├── capsule.yaml
├── trace.jsonl
├── model-calls.jsonl
├── tool-calls.jsonl
├── env.lock
├── redaction-proof.json
├── replay.yaml
└── ...

The idea is simple:

The execution should leave behind something you can keep, inspect, compare, replay, sign, and share.

The research paper models the capsule using 15 entity types, including model calls, tool calls, file events, network calls, human approvals, environment information, lineage, redaction records, replay information, and tamper-evidence.

Five verbs explain most of NovaFabric

NovaFabric: from AI agent runs to verifiable evidence - capture, seal, replay, diff, audit

Capture → Seal → Replay → Diff → Audit

1. Capture

First, capture what happened during the execution:

nova capture python my_agent.py

NovaFabric can wrap a command without requiring the application itself to be rewritten. It aims to capture model calls, tool calls, files, network activity, environment, inputs and outputs. The result is a portable Run Capsule.

2. Seal

Capturing information is not enough. If somebody modifies the record afterward, we want that modification to be detectable.

NovaFabric combines established technologies:

  • DSSE for signing
  • RFC 3161 trusted timestamps
  • Merkle logs for append-only history
  • cryptographic redaction records

The important point: NovaFabric is not inventing new cryptography. It combines existing standards into one evidence system for AI-agent executions. That integration, rather than a new signing algorithm, is the paper's main contribution.

Before sealing:  "This file says the agent did X."

After sealing:   "This file says the agent did X,
                  and we can detect if somebody changes it later."

There is an important limitation here. Tamper-evident does not mean tamper-proof. NovaFabric can detect later modification of properly sealed evidence. It cannot prove that a trusted signing party told the truth when the evidence was originally created. The paper explicitly defines this trust boundary.

3. Replay

AI systems have another problem: reproducibility. Run the same prompt tomorrow and you may not get the same answer. So NovaFabric does not pretend every AI execution can be reproduced bit-for-bit. Instead, it defines four replay modes:

  • Forensic replay - execute nothing; reconstruct and inspect what happened (incident investigation, security review, audit, debugging).
  • Mocked replay - use previously recorded model responses instead of calling the model again, to test the surrounding application logic.
  • Exact replay - try to recreate the same environment, model, data, and execution, when deterministic conditions exist.
  • Semantic replay - run against another model and compare the meaning of the new output with the original, e.g. when upgrading from Model v1 to v2.

The paper is deliberately careful here: only the mocked mode is deterministic by construction with respect to model inputs, while forensic mode executes nothing.

4. Diff

Run B produces a bad result. What changed? Maybe the model, the prompt, a dependency, a tool response, the environment, or the output. NovaFabric provides a structural diff between capsules, comparing the structure of runs instead of giant text logs.

In the paper's controlled experiment, structural diff correctly localized all 140 injected mutations across seven tested mutation classes. That could make questions like this easier to investigate:

"It worked yesterday. Why did the agent fail today?"

5. Audit

Finally, the evidence can be packaged into an Evidence Bundle. Another person should not have to trust your database; you can hand them a portable artifact containing the evidence and the cryptographic information needed to verify it.

Agent execution → Run Capsule → Cryptographic seal → Evidence Bundle → Auditor / Researcher / Security Team

The paper specifies an offline verification path using standard technologies rather than requiring the verifier to trust a NovaFabric server. However, it also states that this independent-verifier interoperability path had not yet been exercised by an independently implemented verifier at publication time. That distinction matters.

Why not just use OpenTelemetry?

I like OpenTelemetry, and NovaFabric uses OpenTelemetry concepts. But the problems are different:

Observability Execution evidence
What is happening? What happened?
Find errors Reconstruct the execution
Metrics and traces Portable run artifact
Operational debugging Audit and forensics
Usually mutable storage Tamper-evident evidence
Live system Historical execution

NovaFabric is not intended to replace Prometheus, OpenTelemetry, Langfuse, LangSmith, or similar systems. It focuses on another question:

Can a past execution be replayed, compared, and proven?

What about secrets?

Capturing everything creates another problem: an agent may see API keys, tokens, passwords, or credentials, and you do not want them inside a bundle sent to another company or auditor. NovaFabric includes secret redaction and a redaction attestation.

But a redaction attestation means "the scanner says it detected and removed this secret." It does not prove that an unknown secret does not exist somewhere else in the capsule. The paper makes this explicit: completeness depends on the scanner, its rules, and the locations it checks. Being clear about these boundaries is especially important for security tools.

What did the research evaluation find?

Model calls captured:                 100% in tested scenarios
MCP tool calls captured:              100% in tested scenarios
Tested tampering classes detected:    3 / 3
Structural mutations localized:       140 / 140
Redaction (14 credential types):      14 / 14 removed
Provenance graph, 10M-edge blast-radius p99:   45.5 ms
Provenance graph, 100M-edge single-client p99: 167.9 ms

But some results exposed weaknesses too. Mocked replay avoided live model calls in 10/10 tested scenarios, but only 2/10 tool-using workloads completed, because tool responses were not yet being substituted in the evaluated version. That was one of the most useful outcomes of the research: the evaluation did not only produce good benchmark numbers, it found real defects.

An unexpected lesson

"The pipeline ran successfully" is not the same as "the evidence is correct."

Several problems were discovered because the tests checked what should have been recorded, rather than simply checking whether the program exited successfully. At one point some evidence streams were missing even though capture appeared to work. In another experiment, an ingest service returned successful HTTP responses even though its configured metadata database had no tables.

"Did it run?" is a weak test.
"Can we prove the expected things actually happened?" is a much stronger one.

NovaFabric today

NovaFabric is open source and self-hosted. The project has moved beyond the exact version evaluated in the paper, so the paper should be read as a research snapshot, not a complete description of today's repository.

pip install novafabric

nova capture python my_agent.py
nova validate <capsule>
nova replay <capsule> --mode forensic
nova diff <capsule-a> <capsule-b>

The project is still beta. Some local functionality is more mature, while server, cluster-scale, and some at-scale components remain experimental. I do not want to describe experimental infrastructure as production-proven.

The bigger idea

As AI agents move from chat interfaces into systems that actually do things - changing production infrastructure, approving financial operations, modifying software, moving important data, operating HPC systems, coordinating other agents - observability alone will not be enough. Eventually someone will ask: What exactly happened? Can you prove it? Can you reproduce it?

That is the space NovaFabric is exploring. Not another agent framework, not another tracing dashboard, but infrastructure for turning an execution into something we can capture → seal → replay → compare → audit.

Try it

NovaFabric is Apache-2.0 and open source:

If you work on AI agents, MLOps, HPC, security, observability, provenance, or reproducibility, I would love your feedback. Especially this question:

When your AI agent makes an important decision today, what evidence will you still have six months from now?

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.