Dev.to AI 🤖 Ai 👁 0 📖 4 min read

How I Stopped 2 AM Outage Loops Using Hindsight Memory

You know the feeling. Pager goes off at 2 AM. You stumble to your laptop, open the terminal, and stare at a PostgreSQL stack trace you've seen before — but can't remember exactly what fixed it last time. You dig through

You know the feeling. Pager goes off at 2 AM. You stumble to your laptop, open the terminal, and stare at a PostgreSQL stack trace you've seen before — but can't remember exactly what fixed it last time. You dig through Slack threads, closed Jira tickets, and a half-finished runbook. Forty minutes later, you find the command. The outage ends.

Three months pass. The same error fires again.

That cycle ends today.

The Problem: SRE Amnesia

Traditional incident response relies on humans remembering patterns across outages. Post-mortems get written, filed in Confluence, and forgotten. The institutional knowledge lives in the heads of senior engineers — until they change teams.

What if your AI agent could remember every verified fix, forever, and surface the right one the instant a matching failure is detected?

That's exactly what I built with KubeHeal AI using Hindsight — a lightweight, open-source agent memory engine from Vectorize.

Architecture: The Dual-Action Cycle

KubeHeal operates on a simple three-phase loop:


RECALL → DIAGNOSE → RETAIN

  1. RECALL — On every incoming incident, query Hindsight for semantically similar past resolutions
  2. DIAGNOSE — Inject recalled context into Groq (Qwen3-32B) for an authoritative, cluster-specific fix
  3. RETAIN — Once resolved, write the verified post-mortem back into Hindsight

Every outage makes the system smarter. Future on-call engineers receive instant, battle-tested prescriptions instead of generic documentation.

The Stack

Component Purpose
Hindsight Persistent semantic memory — recall and retain
Groq (Qwen3-32B) Fast LLM inference for diagnosis
Python + Rich Terminal UI
python-dotenv Configuration

Core Code: gent.py

`python
import os
from dotenv import load_dotenv
from groq import Groq
from hindsight_client import Hindsight
from rich.console import Console
from rich.panel import Panel

load_dotenv()
console = Console()

hindsight = Hindsight(
base_url=os.getenv("HINDSIGHT_BASE_URL"),
api_key=os.getenv("HINDSIGHT_API_KEY"),
)
groq_client = Groq(api_key=os.getenv("GROQ_API_KEY"))

BANK_ID = "kubeheal-cluster-memory"

def triage_incident(service_name: str, error_log: str) -> str:
# Step 1: RECALL
memories = hindsight.recall(
bank_id=BANK_ID,
query=f"Service: {service_name} Error: {error_log}",
)

# Step 2: DIAGNOSE
prompt = f"""You are KubeHeal, an autonomous SRE.

Historical Post-Mortems (Hindsight Memory):
{memories or "No prior incidents on record."}

Active Incident:
Service: {service_name}
Log: {error_log}

Prescribe the exact verified fix or initial diagnostics. Format: [Diagnosis], [Action], [Prevention].
"""
response = groq_client.chat.completions.create(
model="qwen/qwen3-32b",
messages=[{"role": "user", "content": prompt}],
temperature=0.1,
)
return response.choices[0].message.content

def retain_post_mortem(service_name: str, error_pattern: str, verified_fix: str):
# Step 3: RETAIN
hindsight.retain(
bank_id=BANK_ID,
content=f"Service: {service_name}\nSignature: {error_pattern}\nFix: {verified_fix}",
context=f"cluster_postmortem_{service_name}",
)
`

The Before vs. After Demo

Scenario 1 — Cold Start (No Memory)

The agent hits a PostgreSQL connection exhaustion error for the first time:


FATAL: remaining connection slots are reserved for
non-replication superuser connections (error 53300)

Hindsight returns nothing. Groq correctly says "No historical match found" and suggests generic diagnostics:


[Diagnosis] PostgreSQL max_connections limit reached
[Action] SELECT count(*) FROM pg_stat_activity;
SHOW max_connections;
[Prevention] Tune max_connections or add PgBouncer

The SRE investigates, finds the rogue nalytics-worker container, applies the fix, and retains it:

python
retain_post_mortem(
service_name="billing-worker-service",
error_pattern="PostgreSQL connection slots reserved error 53300",
verified_fix="docker stop analytics-worker && psql -U postgres -c 'SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = ''idle'';'"
)

Scenario 2 — Memory Active ✨

Three weeks later, the same failure fires — phrased differently:


Active DB connections maxed out; connection refused error 53300.

This time, Hindsight instantly surfaces the stored post-mortem. The agent's response changes dramatically:


[Diagnosis] This matches the billing-worker-service incident (error 53300).
Root cause: analytics-worker container leaking idle connections.
[Action] docker stop analytics-worker &&
psql -U postgres -c 'SELECT pg_terminate_backend(pid)
FROM pg_stat_activity WHERE state = ''idle'';'
[Prevention] Add connection timeout to analytics-worker config.
Consider PgBouncer for connection pooling.

Zero investigation. Exact command. Outage resolved in seconds.

Why Hindsight?

Most memory solutions for LLM agents require you to build your own vector database, chunking pipeline, and embedding logic. Hindsight collapses all of that into two method calls:

python
hindsight.recall(bank_id=BANK_ID, query=...) # semantic search over history
hindsight.retain(bank_id=BANK_ID, content=...) # persist a new memory

The agent memory model is simple: every verified resolution is a first-class memory that can be semantically retrieved in future sessions. No fine-tuning, no prompt engineering tricks.

Get Started in 5 Minutes

ash
git clone https://github.com/your-username/kubeheal-ai.git
cd kubeheal-ai
python -m venv venv && venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # add your Hindsight + Groq keys
python demo.py

Get your Hindsight API key at ui.hindsight.vectorize.io —

Resources

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.