CodeSmith How Ledgers and Snapshots Cure AI Debt
Ledger and Snapshots: The Debt AI Owes and the Regret Medicine Left for You Source version of CodeSmith: v0.5.0 (commit 3a74c82f). All paths are relative to the repo root; line numbers refer to this version. Intended
Ledger and Snapshots: The Debt AI Owes and the Regret Medicine Left for You
Source version of CodeSmith:
v0.5.0(commit3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: readers who have inherited a module that AI refactored, or who in the middle of the night wanted to roll back but had no idea which step to roll back to.
Start with two moments every user of AI coding tools knows intimately.
Moment one: three weeks later, you open a module that AI refactored. Inside there is a shim whose comment reads "temporary compatibility, remove later" β it has been three weeks; there are three functions that a search of the entire repository cannot find a single caller for; and there is a helpers.rs whose functionality overlaps heavily with the utils.rs you wrote three years ago. Who left these? Why? Nobody knows. Worse: the next Agent walks in, glances around, accepts the whole lot as established architecture, and keeps building on that foundation.
Moment two: late at night, the Agent announces, brimming with confidence, "Done β all tests pass." You run git status: nineteen files changed, six of which have nothing to do with what you originally asked for. You want to roll back β but to which step? You never had the habit of committing at every step, and within a single step the Agent may have invoked ten tools.
For each of these two moments CodeSmith has a mechanism ready: a ledger (the Slop Ledger) and a dose of regret medicine (side-git snapshots). Of everything in this project, these are the two designs I most want to press on my peers β because almost nobody is doing serious work in either direction.
1. The Slop Ledger: Registering AI's Technical Debt
The module documentation's very first paragraph puts the bullseye on display (crates/agent-runtime/src/slop_ledger.rs:1-9):
AI agents often leave behind invisible "slop" after a task: compatibility shims, unmigrated callers, duplicated concepts, naming drift, stale docs/tests, suspected dead code, and tool gaps.
The Slop Ledger makes this residue visible and queryable so the next agent (or human) doesn't rediscover it, amplify it, or mistake it for intended architecture.
Note the final infinitive: mistake it for intended architecture β taking it for architecture that was deliberate. That is the root of moment one's disease: AI's residue carries no signature, no provenance, no bookkeeping, so the next intelligence β human or machine β has no way to tell "this is considered design" from "this is scrap from the last rushed job". And so the scrap gets inherited, gets amplified, and eventually fossilizes.
The word "slop" is also a choice both mean and precise. The industry's term of art is debt β a financial metaphor, implying interest and a repayment schedule. But what AI leaves behind often does not even qualify as "debt": nobody intends to repay it, and sometimes nobody even knows it exists. It is slop β literally swill, the drippings of the assembly line β and the only real problem is that nobody mops it up.
The Skeleton of the Ledger
The ledger is a JSON file (~/.codesmith/slop_ledger.json); each record is a SlopEntry (slop_ledger.rs:179-209) whose fields include the classification bucket, severity, confidence, owner (a person, a team, or "auto"), source links (file path/line number/URL), a one-line title, a detailed description, lifecycle status, a cleanup recommendation, timestamps, and an optional linked task ID.
Classification is ten buckets (SlopBucket, slop_ledger.rs:39-50):
pub enum SlopBucket {
RetainedCompatibility, // retained compatibility layer
UnmigratedCallers, // callers not fully migrated
DuplicateConcepts, // duplicate concepts
NamingDrift, // naming drift
StaleDocs, // stale docs
StaleTests, // stale tests
SuspectedDeadCode, // suspected dead code
UnverifiedPublicBehavior,// unverified public behavior
ToolGaps, // tool gaps
AcceptedDebt, // accepted debt
}
Three of the buckets deserve to be savored on their own. SuspectedDeadCode β note the word "suspected", working with the standalone confidence field: the model's judgment that "nobody uses this code" is often wrong, so an entry must carry its uncertainty with it; that is the brake reserved for high-risk cleanup moves like "delete the dead code". ToolGaps β when the Agent, mid-job, discovers "I want to do X but there is no tool for it", that too is slop: a debt owed by the workflow. AcceptedDebt β the bucket with the best sense of humor: an admission that some debts are simply written off. Not every residue needs cleaning; only when "we know about it, and we have decided to live with it" can also be booked does the ledger become complete β otherwise it is just another anxiety generator.
The ledger comes with four tools: slop_ledger_append (book an entry), slop_ledger_query (query), slop_ledger_update (update status), slop_ledger_export (export). The export product is a redacted Markdown handoff document, which the module docs say outright is "suitable as a GitHub issue, or as a compaction relay" (slop_ledger.rs:19-20) β the latter meaning the ledger can survive a /compact context compaction and go on serving as the session's long-term memory.
Why This Is a One-of-a-Kind Design
Every AI coding tool on the market is optimizing the process of doing the work β faster, more accurate, cheaper (which is what Article 4 of this series was about, too). But almost no tool seriously handles the residue after the work is done. In human-run engineering this job is held up by code review, by senior engineers' memory, and by the oral tradition of "who on earth would dare touch this". In the AI era all three collapse: an Agent has no memory across sessions, and the next Agent is a total stranger.
The Constitution's Article VI had this case on file long ago (crates/agent-runtime/src/prompts/base.md:37-39):
The next intelligence β human or machine β should not have to re-discover what you already learned.
The Slop Ledger is that clause rendered as accounting: it turns "what the next intelligence should not have to re-discover" from oral tradition into a queryable data structure. In fairness: the Constitution itself does not force a ledger entry every turn (this is tool surface, not a command), but a "ledger left for the next intelligence" and a "tidy workspace left for the next intelligence" are plainly the same worldview implemented twice.
2. side-git Snapshots: Regret Medicine Every Turn
The second mechanism deals with moment two. The module docs open (crates/agent-runtime/src/snapshot/mod.rs:1-13):
Each turn the engine takes a
pre-turn:<seq>snapshot of the user's workspace into a side git repo at~/.codesmith/snapshots/<project_hash>/<worktree_hash>/.git, then a matchingpost-turn:<seq>snapshot when the turn finishes. Users can roll back via/restore N(slash command) or, when the model recognises an "undo my last edit" intent, therevert_turntool.
How low does the operating cost of a rollback go? You do not even have to remember the command β say to the Agent, "No, undo that change just now," and it calls revert_turn itself.
Why Open a Second git Off to the Side?
The keyword of the design is side (a bypass lane). The module docs give three reasons (snapshot/mod.rs:12-24):
-
The user's
.gitis never touched. Every git invocation sets--git-dirand--work-treeas a pair β the docs make a point of stressing that "that single invariant is what keeps snapshots and the user's repo completely independent". Your branches, your staging area, your reflog: the snapshot system does not lay a finger on any of them. - Workspaces without git still get snapshots. A snapshot is not a privilege reserved for git repositories.
- git's own deduplication keeps the disk under control. Content-addressed storage shrinks a "100 MB workspace Γ 12 turns" scenario by 10-30Γ.
The density of engineering detail is where a "plain feature" like this shows its class: 7-day retention by default (pruned at session start), gc.auto = 0 on the side repo β so that background gc never suddenly opens fire mid-turn β followed by an explicit git gc --prune=now after the prune; tmp_pack_* temp files left behind by interrupted pack operations get a dedicated startup sweep; at most 50 snapshots per workspace, oldest deleted first once the cap is exceeded (snapshot/mod.rs:44-46, #1112). Behind each of these small decisions is a real crash.
The Failure Model: A Safety Net, Not a Gate
But what gives this module the most personality is its failure model (snapshot/mod.rs:29-37):
Pre/post-turn snapshot calls are non-fatal. If
gitis missing, the disk is full, or the workspace is on a read-only filesystem, the turn proceeds and the engine logs a warning. The snapshot is a safety net, not a correctness gate.
Put that sentence next to Article 4's capacity controller, which was sentenced to "off by default", and you can see a consistent sense of order: the capacity controller rewrites history and injures the cache, so it must stay off however exquisitely it is built; the snapshot system, on failure, simply absents itself β it would rather not be there than stand in the user's way. Two sets of trade-offs pointing in opposite directions, one principle: an auxiliary mechanism does not get to be a roadblock. In a harness, every component must know exactly which class of citizen it is β the main loop is the first-class citizen under the Constitution's protection, and the rest are all there to serve.
The Calculus of the pre/post Double Snapshot
Why two snapshots per turn instead of one? pre-turn is the insurance; post-turn is the evidence. When something goes wrong you have two ways back: to before a given turn (treat that turn as never having happened) or to after it (keep that turn; what you undo is the turns that followed). A single snapshot can only answer "start over from the top"; only a double snapshot can answer "return precisely to any boundary of your choosing" β with 19 mangled files, the snapshot sequence is the chain of custody showing at which step those 6 unrelated changes began to slip in.
Conclusion: Two Kinds of Honesty About What Has Already Happened
The ledger and the snapshot are really a pair of answers to one question: AI has already done the deed β how do you face it?
The snapshot records state β what the workspace looked like at each turn; it faces the past and serves precise undo. The ledger records semantics β what is garbage, why it is there, who owns it, how to clear it; it faces the future and serves never making the same trip again. One handles "what happened"; the other handles "what it means".
By now this catalog of distrust has three lines: distrust of the output (Article 1, stripping forged calls), distrust of the code (Article 6, pinning the child model), distrust of the cleanup (this installment: ledger and snapshots). Article 1 said that a good harness is a clearly written catalog of distrust β and two more lines are still to come: distrust of judgment (Article 13, business rules written as a goalkeeper) and distrust of claims (Article 15, replaying "tests all green" to reconcile). Flip the catalog to its last page and only one question remains:
In this system, is there anything that trusts you?
Yes. Article III of the Constitution says "the user is the sovereign of this session" β the entire customization system is that trust's construction site, and Article 2 has already taken it apart once. The next four installments, on context, change direction: the harness does not merely restrain the model; it must also govern the model's field of vision.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.