Title: Giving On-Call Engineers a Memory: Building On Call Memory with Hindsight
Article Production incidents have a strange property: the same ones keep coming back. A database connection pool fills up in March, a different service leaks connections in August, and each time an engineer starts from
Article
Production incidents have a strange property: the same ones keep coming back. A database connection pool fills up in March, a different service leaks connections in August, and each time an engineer starts from zero. The knowledge exists somewhere, in a postmortem doc, a Slack thread, or the head of someone who left the company. It just isn't available at 3 a.m. when the pager goes off.
For the Hindsight hackathon, our team built OnCall Memory, an incident response agent for a fictional fintech, Northwind Pay. It analyzes a new alert and stack trace, then recommends a root cause and fix. Its main feature is that it remembers every past incident, what fixed it, and what made things worse.
The problem with stateless assistants
A general-purpose LLM can read a stack trace and suggest that a connection pool is exhausted. That advice is generic. It doesn't know that your team's checkout-service leaks connections from Celery tasks, or that rebooting the RDS instance last spring caused an 85-minute thundering-herd outage. Without memory, an assistant gives the same answer on the twentieth incident as on the first.
Architecture
OnCall Memory is a FastAPI backend with a single-page frontend. When an engineer submits an alert, the backend runs the same analysis twice, side by side:
Without memory: the LLM sees only the alert and logs.
With memory: the backend first queries Hindsight for relevant past incidents and passes them to the LLM as context.
The UI shows both answers next to each other, so the value of memory is visible in seconds. We added a Demo Mode that walks through three scenarios: an initial alert, the same root cause under different symptoms, and a case where the agent warns against a fix that failed before.
How we use Hindsight
Hindsight gives an agent three operations, and we used all three.
Retain. We seeded a memory bank named northwind-oncall with 25 realistic incidents. Each includes a date, affected service, error logs, root cause, fix steps, time to resolve, and whether the fix worked. Hindsight breaks these into smaller memories, so 25 incidents became about 90 entries. Retain is also the learning loop. After an incident, the engineer marks the fix as worked or failed and describes what happened. That text is retained, and the memory counter in the UI goes up. The next similar alert can use it immediately.
Recall. When a new alert arrives, we query the bank with the alert text and logs. Hindsight returns the most relevant memories, which we show in the UI as inspectable chunks with their source incident. Showing them matters. A judge, or a skeptical engineer, can see which past incident drove the recommendation.
Reflect. The Systemic Insights button uses reflect across the whole bank. Instead of matching one alert, it asks what patterns keep causing incidents. In our data it found recurring database connection exhaustion, Redis memory problems from missing TTLs, Kafka consumer rebalance loops, and failures from automated infrastructure changes. It also proposed permanent fixes such as connection proxies, TTL enforcement, and pre-deploy validation.
What memory changes
The difference shows up over several interactions. In the first, the agent gives a reasonable diagnosis backed by similar incidents. In the second, the symptoms differ, but recall links the alert to the same underlying cause. By the third, the agent has learned from a failure and tells the engineer not to reboot the database, because that made things worse last time. The recommendation improves because the agent accumulated experience, not because the prompt got better.
Design choices
Visible memory. Recalled chunks are shown, not hidden, because trust in an incident tool depends on evidence.
Failed fixes count. Remembering what didn't work is as valuable as remembering what did.
Realistic data. Incident IDs, log lines and numbers look like real ones, which makes the demo believable.
Where it could go
A production version could ingest PagerDuty alerts and postmortems automatically, run one memory bank per team, and track which fixes worked over time.
OnCall Memory shows that the gap between a generic assistant and a useful one is often memory. Hindsight made adding it a matter of three well-designed operations.

Code:https://github.com/Divyasri03-ux/oncall-memory
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.