I taught my incident agent to say "I don't remember this"
The most useful thing my agent remembers is what didn't work <!-- AUTHOR: VASUNDHARA SADULA ANGLE: Negative knowledge + human-in-the-loop as the gate into memory TARGET: 1,200-1,500 words Every incident retros
The most useful thing my agent remembers is what didn't work
<!--
AUTHOR: VASUNDHARA SADULA
ANGLE: Negative knowledge + human-in-the-loop as the gate into memory
TARGET: 1,200-1,500 words
Every incident retrospective I have ever read documents the fix. Almost none of them document the
forty minutes spent on the thing that did not work.
That omission is expensive. Our incident history contains four separate connection-pool outages
across three services. In every single one, somebody's first instinct was to raise the pool
ceiling. In every single one, it delayed saturation by ten to fifteen minutes and fixed nothing.
Four times, four different engineers, the same dead end — because nobody wrote it down in a place
the next person would find.
So when I built an incident response agent with persistent memory, I made failed remediations a
first-class thing to remember.
The system, briefly
IncidentMind ingests a production incident, extracts evidence from logs and metrics, queries
organizational memory for similar past incidents, and produces root cause hypotheses that cite
prior experience. Persistent memory comes from
Hindsight.
Two design decisions define the product, and they are related:
- The agent never acts. It recommends, and a human approves.
- Nothing enters memory until a human confirms it.
Remembering the dead ends
There are seventeen memory kinds in the system. The one that earns its place most consistently is
REMEDIATION_FAILURE.
It gets written from three sources. The post-mortem's "what failed" section:
for failure in pm.what_failed[:6]:
add(MemoryKind.REMEDIATION_FAILURE,
f"In {incident.incident_id} on {svc}, this did NOT work: {failure}. "
"Do not spend time on this approach for this failure mode without new evidence.")
Any recommendation the on-call engineer rejected:
for action in incident.recommended_actions:
if action.status == ActionStatus.REJECTED:
add(MemoryKind.REMEDIATION_FAILURE,
f"During {incident.incident_id} on {svc}, the recommendation "
f"'{action.kind.label} on {action.target}' was rejected by the on-call engineer. "
f"{action.rationale}".strip())
And direct engineer feedback with a did_not_resolve verdict, which gets written the moment it is
given rather than waiting for the post-mortem — because "we tried this and it didn't help" is
useful to the next incident even while this one is still burning.
That last one surprised me. I had originally batched all memory writes to the end of the incident
lifecycle, which is tidy but wrong. An incident can run for hours. If a parallel incident hits the
same service in that window, the knowledge should already be available.
Using it
Recalling a failure is only half of it. The agent has to act on it, which happens in two places.
The prompt tells the model explicitly:
HOW TO USE THIS MEMORY:
- If a past incident shares the root cause, say so explicitly and cite its id.
- If a remediation is recorded as DID NOT WORK for this failure mode, do NOT
recommend it again; if you must, justify why this time is different.
- Prefer a runbook this organization has already used successfully.
But I do not rely on the model honouring that. Recommendations are cross-referenced against known
failures in code:
def _match_failed_remediation(kind, target, failures) -> str | None:
needle_words = set(kind.value.split("_")) | set((target or "").lower().split())
needle_words.discard("service")
for text, incident_id in failures.items():
hits = sum(1 for word in needle_words if word and word in text)
if hits >= 2 or kind.label.lower() in text:
return (
f"Organizational memory: this was tried during {incident_id} and did not "
f"resolve the incident. Confirm why this time is different before approving."
)
return None
A matching action gets a warning banner and is sorted to the bottom of the list — but it is not
hidden. That distinction mattered more than I expected when I watched people use it. Sometimes
raising the pool ceiling genuinely is right, when the cause is a real traffic surge rather than a
leak. Our own history contains exactly that case: a flash sale that quadrupled traffic, where
raising the ceiling was the correct call.
Hiding the option would have taught the agent a superstition. Showing it with the context — this
failed here before, here is why, decide — puts the judgement where it belongs.
[SCREENSHOT: an action card showing the memory warning]
The gate
The second decision is that memory only accepts human-confirmed knowledge.
def learn(self, incident_id: str, engineer: str = "on-call"):
incident = self.get(incident_id)
if incident.postmortem is None or not incident.postmortem.confirmed:
raise InvalidTransition(
"Organizational memory only accepts human-confirmed knowledge. "
"Confirm the post-mortem first."
)
An unconfirmed AI hypothesis never becomes organizational memory. This is not a philosophical
position, it is a corruption-prevention measure. The agent produces hypotheses continuously, most
of which are wrong in some detail. If those flowed into memory, the bank would fill with confident
speculation that later incidents would recall as established fact — and there is no obvious moment
at which anybody would notice.
The workflow makes the gate visible:
Investigate → Resolve → Generate post-mortem → Confirm → Learn from incident
Each step is disabled until its precondition is met. Before committing, the engineer can preview
exactly what will be written:
GET /api/incidents/{id}/learnings/preview
17 durable memories would be written:
[incident_summary] Incident INC-2026-0431 (SEV-2) on the payment-api service...
[root_cause] The confirmed root cause of INC-2026-0431 was a retry wrapper...
[remediation_failure] In INC-2026-0431, this did NOT work: raising the connection pool...
[lesson] Flat request rate with climbing pool waiting means a leak...
Every action is simulated, and I say so everywhere
There is no shell execution path in this codebase. No subprocess, no cloud SDK, no cluster
client. Approving an action produces a realistic transcript and a database state change.
I enforce that structurally rather than by convention, because conventions erode:
def test_no_shell_execution_path_exists_anywhere(self):
forbidden_modules = {"subprocess", "pty", "commands"}
forbidden_builtins = {"eval", "exec"}
forbidden_attributes = {"system", "popen", "Popen", "spawnl", ...}
for path in app_root.rglob("*.py"):
tree = ast.parse(path.read_text(encoding="utf-8"), filename=str(path))
for node in ast.walk(tree):
...
assert not offences, "shell/eval execution path found: " + "; ".join(offences)
My first version of this test grepped the source text and failed immediately — on a docstring where
I had written "no shell, no subprocess, no cloud SDK". Walking the AST fixes that, and also
distinguishes re.compile (fine) from a bare compile() (not fine).
Actions come from a closed vocabulary too. A model that returns "reboot_the_whole_datacenter"
gets coerced into the nearest valid ActionKind rather than passed through:
@field_validator("kind", mode="before")
@classmethod
def _coerce_kind(cls, v: Any) -> Any:
raw = str(v or "").strip().lower().replace(" ", "_").replace("-", "_")
for member in ActionKind:
if raw == member.value:
return member
aliases = {"restart": ActionKind.RESTART_SERVICE, "rollback": ...}
...
The model cannot introduce an action the system has no simulation for. That is a much stronger
guarantee than trusting a prompt.
What it looks like when it works
Same service, three weeks apart. The first incident has nothing in memory, and the agent says so:
"No relevant organizational memory found. Proceeding with first-principles investigation." Low
confidence, generic hypothesis.
We resolve it, confirm the post-mortem — including the note that raising the pool ceiling bought
twelve minutes and fixed nothing — and the knowledge is retained.
The second incident is written by a different engineer in completely different language. The agent
recalls the first one, explains why, raises its top hypothesis to high confidence citing the prior
confirmed root cause, and attaches a warning to the pool-resize recommendation.
The engineer skips the dead end. That is the entire value proposition, and it comes from a memory
that most incident systems never capture.
Lessons
Record failures deliberately. They will not appear on their own. Add the verdict to the
workflow so capturing it is a click, not an act of discipline.
Warn, do not hide. Context beats suppression. The engineer has information the agent does not.
Write feedback immediately. Incidents overlap. Knowledge batched to the end of a lifecycle is
unavailable exactly when a parallel incident could use it.
Gate writes on human confirmation. An agent that writes its own guesses to memory will
eventually recall them as fact.
Enforce safety structurally. An AST test outlives a code review comment.
The honest limitation: this only works if engineers actually record what failed. The system makes
it easy — four one-click verdicts on every investigation — but it cannot make anyone thoughtful.
In a real deployment I would expect the first few months of memory to be thin, and I would expect
the negative knowledge to lag the positive, because writing down a success feels better than
writing down forty wasted minutes.
Worth it anyway. The four connection-pool incidents in our history cost roughly an hour of
collectively rediscovering the same dead end. One REMEDIATION_FAILURE memory would have paid for
the entire feature.
If you are building agents that learn from experience, the
Hindsight documentation covers the retain and recall model, and
Vectorize's agent memory overview is a good
explanation of why this is a different problem from document retrieval.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.




