OnCall Memory: Building an Incident Response Agent That Learns From Production History
OnCall Memory: Building an Incident Response Agent That Learns From Production History** Production incidents are rarely completely new. A database connection pool can become exhausted again. A deployment configuration
OnCall Memory: Building an Incident Response Agent That Learns From Production History**
Production incidents are rarely completely new.
A database connection pool can become exhausted again. A deployment configuration can break a service again. A payment provider can become rate-limited again. The symptoms may change, but the underlying patterns often repeat.
The problem is that a typical AI assistant starts each incident with very little knowledge of an organization's previous experiences.
It can analyze the logs in front of it and provide technically reasonable suggestions, but it does not automatically know what happened during the last incident, which fix worked, or which attempted solution failed.
That is the problem I wanted to address with OnCall Memory, an incident-response website designed around persistent AI memory.
The idea
OnCall Memory is built for a fictional fintech company called Northwind Pay.
When an engineer receives a production alert, the website provides two perspectives:
Without Memory β a response generated from the current incident.
With Hindsight Memory β a response generated after retrieving relevant historical incidents from the organization's memory.
This side-by-side comparison is the central idea of the website.
The goal isn't simply to make an AI response longer. It is to give the model access to something it normally doesn't have: the team's previous incident experience.
Building the incident history
The project contains a collection of realistic Northwind Pay production incidents.
The incident dataset includes situations such as PostgreSQL connection pool exhaustion, Redis eviction, expired TLS certificates, bad deployment configuration, Kafka consumer lag, OOMKilled pods, and third-party payment API rate limiting.
Each incident contains information such as the symptoms, logs, root cause, resolution steps, outcome, and whether attempted fixes worked.
This is important because a useful memory isn't just:
"There was a database incident."
It should capture the experience surrounding that incident.
What happened?
Why did it happen?
What was tried?
What actually fixed it?
Did the first solution fail?
That information becomes useful context for future incidents.
Hindsight as the memory layer
Hindsight is at the center of the architecture.
The website uses Hindsight to perform three important operations: retain, recall, and reflect.
When useful incident information is available, it can be retained as memory.
When a new incident arrives, the system recalls relevant historical incidents before generating the memory-enhanced diagnosis.
Finally, the website can use reflection to look across the accumulated incident history and ask a broader question:
What patterns keep causing our incidents and what should we fix permanently?
This moves the system beyond responding to individual alerts and toward learning from patterns across incidents.
The architecture
The website uses a lightweight architecture.
The frontend is a single-page HTML, CSS, and JavaScript dashboard.
The backend is implemented with FastAPI and handles the incident analysis workflow, memory operations, and language-model requests.
Groq provides the language-model layer, while Hindsight provides persistent memory.
The project also separates configuration from application code. API credentials are kept locally in environment variables rather than being hardcoded into the source code, while .env.example provides the required configuration structure.
Memory changes the workflow
Suppose a new payment failure arrives.
Without historical context, an AI assistant might suggest checking the payment provider, network connectivity, credentials, rate limits, retries, and application logs.
Those are reasonable suggestions, but the engineer still has to determine which one matches the organization's previous experience.
With Hindsight memory, the workflow becomes different.
The system first looks for similar historical incidents.
If a previous payment incident involved third-party API rate limiting, for example, that historical information can become part of the diagnosis.
The engineer can also see the recalled incidents rather than receiving a completely opaque recommendation.
That visibility was an important part of the design.
Learning doesn't stop after the first answer
OnCall Memory also includes a feedback loop.
After applying a suggested fix, the engineer can indicate whether the fix worked or failed and provide the actual solution.
That information can then become another memory.
The intended loop is:
Incident β Recall β Diagnosis β Fix β Feedback β Retain β Better future diagnosis
This creates a system where future incidents can benefit from experiences that were not available when the original system was built.
The current Hindsight memory bank contains 139 memories.
What I learned
One of the biggest lessons from building the website was that memory quality matters as much as memory quantity.
Adding more memories does not automatically create a better assistant.
A useful incident memory needs meaningful context: symptoms, root cause, attempted fixes, final resolution, and outcome.
I also learned that memory should be visible to the user.
If an AI gives a very specific operational recommendation, an engineer should have some way to understand why that information was relevant.
Displaying recalled incidents makes the memory layer much easier to inspect.
Limitations
OnCall Memory currently uses a fictional Northwind Pay incident history rather than real production data.
That means the system demonstrates the workflow but should not be interpreted as a replacement for an organization's actual incident-management practices.
The quality of the recommendations also depends on the quality of the stored memories. Poor or incomplete incident records can lead to less useful historical context.
Final thoughts
The central idea behind OnCall Memory is simple:
An incident-response assistant should not forget what happened yesterday.
Language models are good at reasoning about the information they receive. Persistent memory adds another dimension: the ability to use accumulated organizational experience.
By combining Groq for reasoning with Hindsight for persistent memory, OnCall Memory turns incident history into a resource that can be recalled during future incidents and analyzed for recurring patterns.
The result is not an AI replacing an on-call engineer.
It is an AI assistant that can increasingly say:
"We've seen something like this before."
ARTICLE 2
From Stateless AI to an AI On-Call Engineer: How Hindsight Gives Incident Response a Memory
Imagine being on call at 2 a.m.
A production alert appears.
You open the logs, start investigating, and ask an AI assistant for help.
The AI gives you a list of possible causes and troubleshooting steps.
It sounds usefulβbut there is a problem.
The AI doesn't know what your team already discovered six months ago.
It doesn't know that a similar incident happened before.
It doesn't know which fix worked.
And it doesn't know which seemingly reasonable fix failed.
That gap is what inspired OnCall Memory, a website for AI-assisted production incident response.
A different approach to AI assistance
Most AI assistants are very good at reasoning about the information included in the current conversation.
But production engineering has another valuable source of information:
experience.
Organizations accumulate years of operational knowledge through incidents.
Someone eventually discovers why a service failed.
Someone tries three different fixes.
One works.
Two don't.
The incident gets resolved, and eventually the details disappear into an incident report, ticket, document, or someone's memory.
OnCall Memory explores what happens when that experience becomes persistent AI memory.
The Northwind Pay scenario
The website simulates incident response for a fictional fintech company called Northwind Pay.
I created realistic historical incidents covering several common production problems, including:
- PostgreSQL connection pool exhaustion
- Redis eviction
- expired TLS certificates
- deployment configuration errors
- Kafka consumer lag
- Kubernetes OOMKilled pods
- third-party payment API rate limiting
The incidents contain realistic operational information including logs, root causes, fixes, outcomes, and failed approaches.
The purpose isn't to create a giant static knowledge base.
It is to create a history that an AI agent can actually use when a new incident appears.
The most important screen: before and after memory
The main interface intentionally shows two answers side by side.
On the left is:
Without Memory
On the right is:
With Hindsight Memory
The first response analyzes the current incident without retrieving organizational history.
The second response first recalls relevant memories from Hindsight and then uses those memories as additional context.
This makes the effect of memory visible.
Instead of simply saying that the agent has memory, the website lets the user observe what happens when memory is introduced into the workflow.
How the memory workflow works
The core workflow can be summarized as:
New Alert β Recall β Analyze β Recommend β Feedback β Retain
When an alert and its logs are submitted, the backend first searches the Hindsight memory bank for relevant historical incidents.
Those results are then used during the analysis.
The resulting response can reference similar incidents, their outcomes, and previous fixes.
After the engineer attempts the recommended solution, feedback can be provided.
If the fix worked, that outcome becomes part of the incident history.
If the fix failed, the failed approach can also be retained.
That distinction matters.
A useful engineering memory should not only remember successful solutions.
It should also remember what not to do.
Hindsight's three roles
Hindsight plays three different roles in the website.
Retain allows new experience to become persistent memory.
Recall allows the agent to retrieve relevant experiences when a new incident occurs.
Reflect allows the system to reason across the accumulated memories and identify larger patterns.
The reflection feature is particularly interesting because it changes the question from:
"How do I fix this incident?"
to:
"Why do these incidents keep happening?"
That second question can lead toward permanent improvements rather than repeatedly fixing symptoms.
A lightweight technical stack
The website deliberately uses a relatively simple architecture.
The backend uses Python and FastAPI.
The frontend is a single-page interface built using HTML, CSS, and JavaScript.
The language-model layer uses Groq.
Hindsight provides the persistent memory layer.
The application explicitly calls the memory operations rather than relying on the language model to decide when memory should be used.
This makes the memory workflow predictable and visible in the application architecture.
Designing for real incident pressure
Incident-response interfaces should not make engineers work harder.
The website therefore focuses on a simple operations-dashboard style.
An engineer can submit an incident, see the two responses, inspect recalled historical incidents, and provide feedback.
The memory panel makes it clear that the answer isn't based solely on the current alert.
There is also a demo flow designed to show how the system behaves across multiple incidents and changing symptoms.
The idea is to demonstrate that memory should influence future behavior rather than being a decorative feature.
The unexpected lesson: failed fixes are valuable
One of the most interesting aspects of the project is the feedback loop.
It is tempting to think of AI memory as a collection of successful answers.
But production engineering doesn't work that way.
A failed troubleshooting attempt can be extremely valuable.
If an engineer tried a particular configuration change during an incident and it didn't solve the problem, remembering that outcome can prevent the same approach from being repeatedly suggested.
That leads to a broader definition of useful AI memory:
Memory is not just what worked. Memory is what the system learned.
What still needs improvement
The current project is a prototype using fictional Northwind Pay data.
A production deployment would need significantly more validation around memory quality, access control, data privacy, observability, and integration with real incident-management systems.
Another important consideration is that retrieved historical incidents should be treated as context, not unquestionable truth.
Past incidents can be incomplete or different from the current situation.
The engineer still needs to validate the diagnosis.
Conclusion
OnCall Memory explores a simple shift in how we think about AI assistants.
A stateless assistant asks:
"What can I infer from this incident?"
A memory-enabled assistant can additionally ask:
"What have we experienced before that might help here?"
That difference can turn organizational incident history into an active part of the troubleshooting process.
With Hindsight providing retain, recall, and reflection capabilities, the website demonstrates how an incident-response agent can accumulate experience rather than starting from zero every time.
The ultimate goal isn't an AI that takes control of production incidents.
It's an assistant that becomes more useful because it remembers.
β οΈ Important before you paste either article
The submission guide says each team member submits one article, so don't have both people submit the exact same article.
Use:
- Article 1 β Member 1
- Article 2 β Member 2
Also, the guide says do not mention the word "hackathon" in the article title, body, or hashtags, so I've deliberately kept it out.
And for the final versions, add 2β4 screenshots of your actual websiteβespecially the Without Memory vs With Hindsight screen and the memory panel.
code:https://github.com/Divyasri03-ux/oncall-memory

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.