Dev.to AI 🤖 Ai 👁 1 📖 6 min read

What coldstart's first users actually broke

Testing your own tool tells you it works. Real usage tells you how it fails — and those turned out to be almost entirely different lists. I've been shipping coldstart in the open for a few months now, and the gap betwee

Testing your own tool tells you it works. Real usage tells you how it fails — and those turned out to be almost entirely different lists.

I've been shipping coldstart in the open for a few months now, and the gap between "works when I use it" and "works when a stranger installs it at 11pm on a machine I've never seen" turned out to be the most useful thing I learned. None of what follows is hypothetical — it's pulled straight from the release history. Real error messages, real root causes, real fixes.

The first bug was the install itself

Before anyone could tell me the tool was useful or useless, most of them had to get past npx coldstart-mcp init actually finishing. It didn't, for a while.

The symptom was a hang at effectively 100% CPU for minutes, sometimes indefinitely. The cause was almost comic once I found it: two tree-sitter grammar packages declared conflicting peer-dependency ranges (^0.21.x vs ^0.22.x), and npm's resolver, given no lockfile to anchor against on a cold cache, just... spun. I found over 89,000 repeated placeDep ROOT lines in the debug log without convergence.

The first fix was a flag — --legacy-peer-deps — and clearer docs. That held for a while, then broke again in a different way: __filename is not defined, a straight ESM/CommonJS mismatch that crashed init before it could do anything, because the tests exercised the module directly and never touched the actual install-path code that broke.

The fix that actually stuck was more drastic: moving the whole parsing engine to WASM in v2.1.0. No native node-gyp compile, no grammar peer-dependency ranges to reconcile, no install scripts. npm install -g coldstart became a completely ordinary install. That's the pattern that shows up everywhere in this list — the first fix addresses the symptom, the real fix removes the category of failure.

A timeout nobody could reproduce

For a while, people reported MCP server "coldstart" connection timed out after 30000ms and I couldn't get it to happen locally no matter what I tried.

The cause was invisible unless you went looking for it specifically: .mcp.json launched the server via npx -y coldstart-mcp, and every single startup, npx re-hashed the entire install tree — fifteen tree-sitter packages, several with multi-megabyte native binaries — as part of npm's supply-chain integrity check. That took 25–30 seconds. The MCP client's timeout was 30. It wasn't flaky; it was a race that depended on exactly how fast someone's disk and CPU were, which is why it never reproduced on my machine and always did on theirs.

The fix wasn't a flag, it was recognizing npx was the wrong tool for a "launch on every editor session" pattern — it's built for one-shot scaffolders, not long-running servers. init now does the expensive install once, then writes a direct node /path/to/index.js invocation. Sub-second startup, no npm involvement at all after setup. Existing users got an auto-migration that rewrote their config on first launch with v1.4.0, timestamped backup included, no action required.

The notebook that forgot things mid-session

This one only showed up in my own long sessions, which is a category of bug that's genuinely hard to catch any other way.

The capture pipeline sliced an agent's transcript by a stored line offset to know what was new since the last capture. /compact in Claude Code rewrites the transcript to something much shorter — and the stored offset kept pointing past the new, shorter end. Every turn after a compaction was silently invisible to capture for the rest of the session. Not an error, not a warning — just quietly nothing, in exactly the sessions long and complicated enough to be the ones most worth remembering.

The fix was a shrink guard: detect when the transcript got shorter than the stored offset, reset it, reprocess from the new baseline. Small fix, but it only exists because I was using the tool on real, messy, long-running tasks — not the kind of thing a unit test would ever surface, since unit tests don't run long enough to hit /compact in the first place.

Recall was mostly noise, and I measured exactly how much

This is the one I'm most glad I didn't just patch on vibes.

The complaint, from watching my own sessions, was that recall — the feature that surfaces a past note automatically when it might be relevant — felt like it was injecting the wrong thing most of the time. Feeling like isn't a fix criterion, so I hand-labeled 140 real injections against whether they were actually relevant. Baseline precision: 31%.

Three separate causes, each measurable on its own:

  • Single ordinary-word matches were mostly homonyms, not topic matches. "Merge the PR" pulled up a notebook-merge note. "Status 403" pulled up an unrelated note just called "status." These were 61% of all injections, at 5% precision. The fix wasn't a score threshold or a stopword list — I tried both, both failed on held-out validation — it was requiring a second corroborating signal unless the term was code-shaped or explicitly named.
  • The same note re-surfaced repeatedly in one session — worst case, twelve times — pure tax, since the content was already sitting in context from the first injection.
  • Harness telemetry was leaking into the relevance query. Boilerplate wrapper text like "while running local commands" was enough on its own to promote an unrelated note two ranks. The stripper for this existed in the code but had never actually shipped.

After all three fixes: 47% precision, keeping 86% of the injections that were actually good. I published the honest caveat alongside it — the numbers come from one repo's notebook and one person's prompt style, and while the rule held on a held-out split, I'm not claiming it generalizes to every codebase. It's real progress, measured, not a marketing number.

Feedback that wasn't a bug report at all

Not everything came from something breaking. The rename from coldstart-mcp to coldstart in v2.0.0 came from a slower-burning realization: the old package metadata — description, keywords — never contained the word "memory" at all, so nobody searching for what the tool actually does could find it by searching for it. That's the same terminology-mismatch problem I've written about elsewhere with SEO, showing up first inside the package registry, before it showed up in search results.

That fix cascaded into its own small comedy of validation errors: the MCP registry rejected the first server.json for a 215-character description against a 100-character limit; npm truncated a separate description mid-word; and one release later, publishing failed outright with a 403 because the registry grants publish rights using a GitHub username's exact casing, and the config had it lowercased. None of these were interesting bugs. All of them were the unglamorous, necessary cost of actually being findable.

What this actually taught me

None of this was found by writing more tests upfront, and I don't think it could have been. Some of it — the compact bug, the recall noise — only exists in the texture of a long, real session, doing a real task, with the kind of context a synthetic test case doesn't have a reason to construct. The rest — the npx timeout, the registry casing — only exists once you're dealing with infrastructure you don't fully control: npm's integrity checks, a registry's authorization rules, someone else's machine and someone else's cold cache.

The pattern, if there is one: almost every fix that actually held up removed a category of failure rather than patching the specific case in front of me. A flag instead of a WASM rewrite fixed the symptom for a week. The rewrite fixed it permanently. That's turned into something close to a rule for how I ship changes to this now — the local, contained fix is often correct as a stopgap, but I don't consider anything actually finished until I can name the class of problem it solved, not just the instance.

coldstart is open source, MIT — install with npm install -g @cstart/coldstart, then coldstart init in any repo. If you hit something that isn't on this list, the GitHub Discussions are open, and the full, unedited history of everything above is in the release notes.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.