How I Stopped 2 AM Outage Loops Using Hindsight Memory
You know the feeling. Pager goes off at 2 AM. You stumble to your laptop, open the terminal, and stare at a PostgreSQL stack trace you've seen before — but can't remember exactly what fixed it last time. You dig through
You know the feeling. Pager goes off at 2 AM. You stumble to your laptop, open the terminal, and stare at a PostgreSQL stack trace you've seen before — but can't remember exactly what fixed it last time. You dig through Slack threads, closed Jira tickets, and a half-finished runbook. Forty minutes later, you find the command. The outage ends.
Three months pass. The same error fires again.
That cycle ends today.
The Problem: SRE Amnesia
Traditional incident response relies on humans remembering patterns across outages. Post-mortems get written, filed in Confluence, and forgotten. The institutional knowledge lives in the heads of senior engineers — until they change teams.
What if your AI agent could remember every verified fix, forever, and surface the right one the instant a matching failure is detected?
That's exactly what I built with KubeHeal AI using Hindsight — a lightweight, open-source agent memory engine from Vectorize.
Architecture: The Dual-Action Cycle
KubeHeal operates on a simple three-phase loop:
RECALL → DIAGNOSE → RETAIN
- RECALL — On every incoming incident, query Hindsight for semantically similar past resolutions
- DIAGNOSE — Inject recalled context into Groq (Qwen3-32B) for an authoritative, cluster-specific fix
- RETAIN — Once resolved, write the verified post-mortem back into Hindsight
Every outage makes the system smarter. Future on-call engineers receive instant, battle-tested prescriptions instead of generic documentation.
The Stack
| Component | Purpose |
|---|---|
| Hindsight | Persistent semantic memory — recall and retain |
| Groq (Qwen3-32B) | Fast LLM inference for diagnosis |
| Python + Rich | Terminal UI |
| python-dotenv | Configuration |
Core Code: gent.py
`python
import os
from dotenv import load_dotenv
from groq import Groq
from hindsight_client import Hindsight
from rich.console import Console
from rich.panel import Panel
load_dotenv()
console = Console()
hindsight = Hindsight(
base_url=os.getenv("HINDSIGHT_BASE_URL"),
api_key=os.getenv("HINDSIGHT_API_KEY"),
)
groq_client = Groq(api_key=os.getenv("GROQ_API_KEY"))
BANK_ID = "kubeheal-cluster-memory"
def triage_incident(service_name: str, error_log: str) -> str:
# Step 1: RECALL
memories = hindsight.recall(
bank_id=BANK_ID,
query=f"Service: {service_name} Error: {error_log}",
)
# Step 2: DIAGNOSE
prompt = f"""You are KubeHeal, an autonomous SRE.
Historical Post-Mortems (Hindsight Memory):
{memories or "No prior incidents on record."}
Active Incident:
Service: {service_name}
Log: {error_log}
Prescribe the exact verified fix or initial diagnostics. Format: [Diagnosis], [Action], [Prevention].
"""
response = groq_client.chat.completions.create(
model="qwen/qwen3-32b",
messages=[{"role": "user", "content": prompt}],
temperature=0.1,
)
return response.choices[0].message.content
def retain_post_mortem(service_name: str, error_pattern: str, verified_fix: str):
# Step 3: RETAIN
hindsight.retain(
bank_id=BANK_ID,
content=f"Service: {service_name}\nSignature: {error_pattern}\nFix: {verified_fix}",
context=f"cluster_postmortem_{service_name}",
)
`
The Before vs. After Demo
Scenario 1 — Cold Start (No Memory)
The agent hits a PostgreSQL connection exhaustion error for the first time:
FATAL: remaining connection slots are reserved for
non-replication superuser connections (error 53300)
Hindsight returns nothing. Groq correctly says "No historical match found" and suggests generic diagnostics:
[Diagnosis] PostgreSQL max_connections limit reached
[Action] SELECT count(*) FROM pg_stat_activity;
SHOW max_connections;
[Prevention] Tune max_connections or add PgBouncer
The SRE investigates, finds the rogue nalytics-worker container, applies the fix, and retains it:
python
retain_post_mortem(
service_name="billing-worker-service",
error_pattern="PostgreSQL connection slots reserved error 53300",
verified_fix="docker stop analytics-worker && psql -U postgres -c 'SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = ''idle'';'"
)
Scenario 2 — Memory Active ✨
Three weeks later, the same failure fires — phrased differently:
Active DB connections maxed out; connection refused error 53300.
This time, Hindsight instantly surfaces the stored post-mortem. The agent's response changes dramatically:
[Diagnosis] This matches the billing-worker-service incident (error 53300).
Root cause: analytics-worker container leaking idle connections.
[Action] docker stop analytics-worker &&
psql -U postgres -c 'SELECT pg_terminate_backend(pid)
FROM pg_stat_activity WHERE state = ''idle'';'
[Prevention] Add connection timeout to analytics-worker config.
Consider PgBouncer for connection pooling.
Zero investigation. Exact command. Outage resolved in seconds.
Why Hindsight?
Most memory solutions for LLM agents require you to build your own vector database, chunking pipeline, and embedding logic. Hindsight collapses all of that into two method calls:
python
hindsight.recall(bank_id=BANK_ID, query=...) # semantic search over history
hindsight.retain(bank_id=BANK_ID, content=...) # persist a new memory
The agent memory model is simple: every verified resolution is a first-class memory that can be semantically retrieved in future sessions. No fine-tuning, no prompt engineering tricks.
Get Started in 5 Minutes
ash
git clone https://github.com/your-username/kubeheal-ai.git
cd kubeheal-ai
python -m venv venv && venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # add your Hindsight + Groq keys
python demo.py
Get your Hindsight API key at ui.hindsight.vectorize.io —
Resources
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.