Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 8 min read

Time-Travel Debugging for LLM Agents: Burr's Counterfactual Replay Architecture

Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. Yo

Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.

The alternative is counterfactual replay: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. When you replay a branch you didn't change, every output hash should come back identical and cost zero tokens. That's the difference between a diff you can trust and a diff that's just noise.

This is how Rewind works on top of Burr, an open-source state machine framework for agent pipelines. The implementation exposes four design rules that make determinism provable and a handful of traps that break it.

The Problem with Traditional Agent Observability

Standard logging captures what happened. It doesn't capture why, and it doesn't let you test what would have happened if you changed one decision upstream.

What you get from a trace tool:

  • A waterfall of step names, timestamps, and token counts.
  • A JSON blob of inputs and outputs per step.
  • No causal relationship between step N and step N+3.

What you can't do:

  • Rewind to step three, change the retrieval query, and see how it ripples through steps four through eight.
  • Prove that your prompt change (not model temperature drift) caused the output shift.
  • Replay a branch without burning tokens on unchanged LLM calls.

This matters in production when you need to debug a specific run that failed, not reproduce a failure class across ten new runs.

Architecture: State Snapshots, Content Hashes, and Replay Logic

The Rewind implementation sits on top of Burr, which already provides a state machine abstraction for agent pipelines. The key additions are persistent state snapshots, per-node telemetry with content hashing, and a replay engine that knows when to use cached outputs.

System Shape

browser (Vite + React 18 + TS + Tailwind)
  โ”œโ”€ Graph ยท NodeEditor ยท DiffPanel ยท CostBar ยท RawJSON ยท Settings
  โ”‚
  โ””โ”€ /api (Vite proxy)
      โ†“
    FastAPI (spawns daemon thread per run; client polls /graph)
      โ†“
    Burr application
      plan โ†’ research โ†’ analyse โ†’ critique โ†’ revise โ†’ compose โ†’ verify โ†’ publish
      (llm)   (tool)     (tool)     (llm)     (llm)     (llm)     (tool)   (tool)
      โ”‚
      โ”œโ”€ SQLitePersister("burr_state") โ† state after every node
      โ””โ”€ NodeTelemetryHook(PostRunStepHook) โ† inputs/outputs/latency/tokens/hash
          โ†“
        rewind.db (SQLite, WAL)
          tables: runs ยท nodes ยท edges ยท cache ยท tool_cache ยท burr_state

Burr's role:

  • Defines the state machine (nodes, edges, transitions).
  • Persists state after every node execution via SQLitePersister.
  • Exposes hooks for telemetry capture.

Rewind's additions:

  • Content hashing of every node's inputs and outputs.
  • A cache table keyed by (node_name, input_hash) that stores (output_hash, output_blob, token_count).
  • Replay logic that checks the cache before invoking the node function.

Four Design Rules for Provable Determinism

Rule Why It Matters Implementation Detail
1. Hash inputs, not timestamps Timestamps always change; you need semantic equivalence. Hash the serialized input dict after stripping metadata keys like timestamp, run_id.
2. Separate tool cache from LLM cache Tool calls can be deterministic (database query) or non-deterministic (API with rate limits). Store tool outputs in tool_cache with a TTL or version tag; LLM outputs in cache with no TTL.
3. Replay only downstream nodes Unchanged branches should return cached outputs without re-execution. Walk the DAG from the fork point; for each node, check if inputs changed. If not, return cached output.
4. Diff by content hash, not text LLM outputs can have whitespace or formatting jitter that doesn't matter. Store sha256(canonical_json(output)) alongside the raw output. Diff hashes first, then show text diff only if hashes differ.

Replay Flow

  1. User selects a completed run and a fork point (e.g., step three: "research").
  2. User edits the input to that step (e.g., changes the retrieval query).
  3. Rewind computes the new input hash and checks the cache.
  4. If cache miss, it invokes the node function and stores the result.
  5. It walks downstream nodes. For each:
    • Compute input hash from upstream outputs.
    • Check cache.
    • If hit and input hash matches, use cached output (zero tokens).
    • If miss, invoke and cache.
  6. Return the new trajectory and diff it against the original.

Cost tracking:

  • Original run: sum of all node token counts.
  • Replay run: sum of only cache-miss nodes.
  • Diff panel shows token delta and highlights which nodes re-executed.

Code: Replay Logic with Cache Lookup

# Simplified replay engine (actual implementation in FastAPI route)

def replay_from_fork(run_id: str, fork_node: str, new_input: dict):
    original_run = load_run(run_id)
    state = load_state_at_node(run_id, fork_node)

    # Start from fork point with new input
    state[fork_node] = new_input
    input_hash = hash_input(new_input)

    # Check cache
    cached = query_cache(fork_node, input_hash)
    if cached:
        output = cached["output"]
        tokens = 0  # cache hit
    else:
        output = execute_node(fork_node, new_input)
        tokens = output.get("usage", {}).get("total_tokens", 0)
        store_cache(fork_node, input_hash, output, tokens)

    state[fork_node + "_output"] = output

    # Walk downstream nodes
    downstream = get_downstream_nodes(fork_node)
    total_tokens = tokens

    for node in downstream:
        node_input = build_input_from_state(node, state)
        input_hash = hash_input(node_input)

        cached = query_cache(node, input_hash)
        if cached and cached["input_hash"] == input_hash:
            # Unchanged branch: use cache
            state[node + "_output"] = cached["output"]
        else:
            # Changed branch: execute
            output = execute_node(node, node_input)
            tokens = output.get("usage", {}).get("total_tokens", 0)
            total_tokens += tokens
            store_cache(node, input_hash, output, tokens)
            state[node + "_output"] = output

    return {
        "run_id": generate_run_id(),
        "forked_from": run_id,
        "fork_node": fork_node,
        "total_tokens": total_tokens,
        "state": state
    }

Key points:

  • hash_input() must be stable: sort dict keys, strip metadata, serialize to canonical JSON.
  • query_cache() returns None if no match or if input hash differs (handles cache invalidation).
  • execute_node() wraps the Burr node function and extracts token usage from the response.

Handling Non-Deterministic Tool Outputs

Not all tool calls are deterministic. API rate limits, timestamp drift, and external state changes break replay.

Strategies:

Tool Type Determinism Cache Strategy
Database query (read-only) Deterministic if schema stable Cache indefinitely, keyed by query hash
External API (weather, stock price) Non-deterministic Cache with TTL (e.g., 5 minutes) or version tag
File write Side effect Don't cache; log the write and replay with a dry-run flag
LLM call Non-deterministic (temperature > 0) Cache by input hash; accept that temperature=0 is required for exact replay

Implementation:

  • Tag each node with a determinism flag: "deterministic", "time-bound", "side-effect".
  • For "time-bound" nodes, store a cached_at timestamp and invalidate after TTL.
  • For "side-effect" nodes, skip caching and log the action instead.

Example:

# In node definition
@action(reads=["query"], writes=["results"], determinism="time-bound", ttl=300)
def fetch_weather(state):
    query = state["query"]
    # API call
    return {"results": call_weather_api(query)}

During replay, if cached_at is older than 300 seconds, re-execute and update cache.

Observability: What You Can See

The Rewind UI exposes:

  • Graph view: DAG with nodes colored by cache hit (green) vs. re-execution (orange).
  • Diff panel: Side-by-side comparison of original vs. replay outputs, with content hash match indicators.
  • Cost bar: Token count for original run, replay run, and delta.
  • Raw JSON: Full state snapshots at each node, with input/output hashes.

Metrics tracked per node:

  • input_hash: SHA-256 of canonical input JSON.
  • output_hash: SHA-256 of canonical output JSON.
  • latency_ms: Wall-clock time for node execution.
  • token_count: Total tokens (prompt + completion) for LLM nodes.
  • cache_hit: Boolean.

Failure modes you can debug:

  • Hash mismatch on unchanged branch: Indicates non-determinism (temperature > 0, external state change, or timestamp leak into input).
  • High token cost on replay: Indicates cache misses due to input drift or cache invalidation.
  • Divergent outputs after fork: Shows causal impact of your input change.

Deployment Shape

Local development:

  • SQLite with WAL mode for concurrent reads during replay.
  • FastAPI runs in a single process; daemon threads handle long-running agent executions.
  • Vite dev server proxies /api to FastAPI.

Production considerations:

  • Replace SQLite with Postgres for multi-instance deployments.
  • Move long-running executions to a task queue (Celery, Temporal) to avoid blocking the API server.
  • Store large outputs (e.g., generated documents) in object storage (S3) and cache only the hash + reference.
  • Add authentication and run isolation (users should only see their own runs).

Scaling:

  • Cache table grows linearly with unique (node_name, input_hash) pairs. Add a retention policy (e.g., delete cache entries older than 30 days).
  • State snapshots grow with run count. Archive completed runs to cold storage after N days.

Traps and Gotchas

1. Timestamp leaks into input hash
If your node input includes a timestamp or run_id, every replay will be a cache miss. Strip metadata before hashing.

2. Non-canonical JSON serialization
Python's json.dumps() doesn't guarantee key order. Use json.dumps(obj, sort_keys=True) or a library like canonicaljson.

3. LLM temperature > 0
Even with identical inputs, the LLM will produce different outputs. Set temperature=0 for deterministic replay, or accept that cache hits only work for exact input matches and you'll need to diff outputs semantically.

4. Tool calls with side effects
Writing a file, sending an email, or updating a database breaks replay. Either skip these nodes during replay (dry-run mode) or log the action without executing.

5. State mutation in node functions
If a node mutates shared state (e.g., a global cache or config object), replay will see the mutated state from the original run. Ensure node functions are pure or reset shared state before replay.

Technical Verdict

Use Burr + counterfactual replay when:

  • You have multi-step agent pipelines with expensive LLM calls.
  • You need to debug specific production runs, not just reproduce failure classes.
  • You want to test prompt changes or tool swaps without re-running the entire pipeline.
  • Your pipeline has a mix of deterministic (database queries) and non-deterministic (LLM calls) steps.

Avoid when:

  • Your pipeline is a single LLM call (no branching to replay).
  • You can't enforce temperature=0 or accept non-deterministic cache misses.
  • Your tools have heavy side effects (payments, external writes) that can't be dry-run.
  • You need real-time replay (cache lookup adds 10-50ms per node).

Alternatives:

  • LangSmith, Weights & Biases: Trace logging without replay. Good for observability, not counterfactual debugging.
  • Temporal workflows: Deterministic replay via event sourcing, but no content-hash caching or zero-token replays.
  • Custom DAG runners (Airflow, Prefect): Task-level retries, but no state-machine abstraction or LLM-aware caching.

The key insight is that deterministic replay requires more than just logging. You need immutable state snapshots, content-addressable caching, and a DAG walker that knows when to skip unchanged branches. Burr provides the state machine primitives; Rewind adds the replay engine and observability layer.

Source Links

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.