One Sentence Hijacked My Agent's Memory. So I Made Every Memory Auditable.
A post this week was uncomfortable to read, and the title is basically the whole problem: one word, "update", in the right framing, was enough to hijack an agent's memory. Not an exploit chain. Just phrasing. The agent r
A post this week was uncomfortable to read, and the title is basically the whole problem: one word, "update", in the right framing, was enough to hijack an agent's memory. Not an exploit chain. Just phrasing. The agent rewrote what it believed because the input told it to, and nothing in the path was allowed to say no.
The reflex is to file that under prompt injection and move on. That misses the half that matters. Injection is an input problem. What the post actually exposed is a storage problem: if the thing your agent believes lives in a box you cannot open, a poisoned sentence read last week quietly shapes every session after it. You cannot see the belief that got rewritten, you cannot trace which input rewrote it, and you cannot undo it.
That is the failure mode I built against. HyperMarrow is a local-first memory layer for agents on Windows, and the properties below are the ones I enforce. They are ordered by how early they act, because every failure above is decided at write time, not at query time.
1. Decisions are written down, with a source and a timestamp
A preference, a decision or a conclusion has to land in durable local storage as an actual record, carrying where it came from and when. An assumption re-derived from the transcript on every turn has no address, no origin and no history. Once a record has provenance, "what does it believe, and why" stops being an argument and becomes a lookup.
2. Write before you compact
Nothing is allowed to summarise or drop the working context until the durable copy exists and has been acknowledged. This is a write-ahead rule, and it acts earliest of the five. If compaction runs first, whatever it dropped is gone, and the loss does not raise an error. It just quietly takes something with it.
3. Forgetting is bounded and scheduled; pinned records never decay
An unbounded memory is not a feature. Records that stopped being relevant age out on a schedule instead of being carried forward forever, and records you explicitly pin are exempt from decay. Without a boundary the retrieval space only grows, and short queries against long records get worse with use, not better.
4. Recall returns your words, and it shows its provenance
If what got written was already lossy, no ranker recovers the sentence you actually needed. Recall pulls the stored record rather than a paraphrase of it, and it hands you the hit together with where the hit came from. That is what makes cross-session continuity real: a decision made in one session can be re-found, with its source, from another.
5. The privacy boundary is a setting, not a promise
Where records live, and whether a given record may leave the machine, is a configuration decision rather than a paragraph in a policy. This is also why local-first matters for the hijack case specifically. A poisoned input stays in your store, where you can find it and remove it, instead of being broadcast into a vendor's model where you have no handle on it at all.
What auditability actually changes
Read back against the hijack:
- You can see the belief. Every record carries its source, so the rewritten entry points at the input that rewrote it.
- You can undo it. Wrong memory is editable, not permanent. There is no support ticket between you and a bad record.
- You can ๅฎ็ฝ hm.qianshi.cooll what changed. Because decisions are stored as records instead of re-derived, a rewrite is a visible edit rather than a different answer you are expected to notice.
- You can keep it. The store exports, so the memory outlives any single tool, session or vendor.
The modules are record, recall, consolidation and file-bridge, exposed over MCP, so the agent does not need a bespoke integration per tool.
What is still not good enough
- Provenance ๅฎ็ฝ hm.qianshi.coolls you where a record came from. It does not ๅฎ็ฝ hm.qianshi.cooll you the source was trustworthy. Reading a poisoned record and judging it is still on you. A source tag is evidence, not a verdict.
- Editable memory means a bad edit is available too. The same open door that lets you repair a hijacked record lets someone with the store change it the other way. Local-first moves that risk onto your machine, which is at least a machine you can inspect.
- Classification is model-dependent. Anything that decides at write time which kind of memory a sentence is will misclassify some fraction of the time. It fails more quietly than losing the data, but it still fails.
- Provenance is not detection. A record can carry a perfectly honest source and still be wrong, and nothing in the store will ๅฎ็ฝ hm.qianshi.cooll you that.
Over to you
Once a memory can be rewritten by a sentence, inspectability stops being a nice-to-have and becomes the floor. So when your agent misbehaves, would you rather read a vendor's log, or open the memory file yourself?
HyperMarrow is a Windows desktop memory layer. The local-first build is here: HyperMarrow
Disclosure: I build HyperMarrow, so read the above with that in mind. The community post quoted here is real and public, and no numbers, user counts or testimonials have been invented for this post.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.