How I Built Proactive Churn Alerts Using Hindsight Reflection
How I Built Proactive Churn Alerts Using Hindsight Reflection How can an AI support copilot detect a customer's growing frustration before they explicitly ask for a manager? Most customer support platforms react to
How I Built Proactive Churn Alerts Using Hindsight Reflection
How can an AI support copilot detect a customer's growing frustration before they explicitly ask for a manager?
Most customer support platforms react to escalation signals too late.
A supervisor is often alerted only after a customer:
- Explicitly demands a manager
- Posts a public complaint
- Threatens a chargeback
- Repeatedly contacts support
By that point, the customer's sentiment may have already deteriorated significantly.
To address this, I designed a support copilot that tracks customer frustration trajectories and surfaces proactive escalation alerts before the customer explicitly demands intervention.
By integrating Hindsight into a FastAPI + React architecture, the system combines historical conversation experiences with Hindsight's reflection engine (areflect) to generate higher-level opinions about customer effort and frustration risk.
This article explains how the system tracks multi-session effort, enforces risk thresholds, and surfaces proactive escalation signals to support representatives.
Architecture Overview
The system operates as a rep-facing support copilot.
It intercepts incoming support queries, retrieves historical customer context, and combines raw experiences with reflection opinions before sending the context to the LLM.
Architecture Diagram
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β React + Vite UI β
β β
β Ticket Queue Conversation Thread Agent Memory Panel β
β Risk Badges Escalation Banner Customer Context β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββ
β
β HTTP / REST API
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FastAPI Backend β
β β
β Ticket Routes Fallback Engine Core Memory β
β /api/tickets Local Cache Pinned Facts β
β β
β Risk & Trajectory Engine β
βββββββββββββββββββββββββ¬βββββββββββββββββββββββββ¬βββββββββββββββββββ
β β
async recall/reflect inference
β β
βΌ βΌ
ββββββββββββββββββββββββββββββββββ βββββββββββββββββββββββββββββββ
β Hindsight Cloud β β Groq LPU Engine β
β β β β
β β’ Retain Experiences β β β’ gpt-oss-120b β
β β’ Scoped Tag Recall β β Primary β
β β’ Reflect & Form Opinions β β β’ qwen3-32b β
β β β Fallback β
ββββββββββββββββββββββββββββββββββ βββββββββββββββββββββββββββββββ
Core Components
| Component | Responsibility |
|---|---|
| FastAPI Backend | API contracts, local caches, trajectory calculation, prompt assembly |
| Groq LPU Engine | Primary and fallback LLM inference |
| Hindsight Memory | Temporal experiences, tagged recall, and reflection opinions |
The architecture uses arecall for historical retrieval and areflect to generate higher-level opinions about customer behavior.
The Core Problem: Customer Effort Happens Over Time
A customer might contact support three times about different minor problems.
Individually, none of those conversations may appear particularly serious.
But together, they represent increasing customer effort.
A stateless LLM looking at only the current ticket cannot easily see this pattern.
For example:
Session 1
Delivery issue
β
Session 2
Replacement issue
β
Session 3
Still waiting
β
Customer effort increases
β
Potential escalation risk
To track this progression, the system combines:
Semantic memory recall + Hindsight reflection + risk invariants
Whenever an issue is marked resolved, an asynchronous reflection task analyzes the customer's historical threads and updates an opinion containing information such as:
- Customer effort count
- Sentiment trajectory
- Repeat-contact pattern
- Overall churn risk
Risk Tier Invariants
One of the most important design decisions was preventing the escalation system from becoming too sensitive.
The system uses three risk tiers:
| Risk Tier | Condition |
|---|---|
| Normal | 1 contact turn + stable sentiment |
| Watch | β₯ 2 contact turns OR declining sentiment |
| Escalate | β₯ 3 contact turns AND declining sentiment |
The critical rule is:
Escalation requires both repeated contact and declining sentiment.
This prevents a single negative message from immediately producing an escalation alert.
Code-Backed Implementation
1. Asynchronous Hindsight Reflection
When a support representative resolves a case, the system triggers Hindsight reflection asynchronously.
async def reflect(
self,
customer_id: str
) -> Optional[dict]:
"""
Triggers Hindsight reflection after an issue
is resolved.
Updates opinions on customer effort trajectory
and frustration risk.
"""
client = self._get_client()
if not client:
return self._local_opinions.get(customer_id)
try:
res = await client.areflect(
bank_id=self.bank_id,
query=(
f"Assess customer effort trajectory, "
f"repeat contact patterns, and "
f"frustration churn risk for {customer_id}"
),
tags=[customer_id]
)
return {
"customer_id": customer_id,
"confidence": 0.94,
"reflection_text": getattr(
res,
"text",
""
) or ""
}
except Exception as e:
logger.error(
f"Hindsight reflect exception: {e}"
)
return None
Why asynchronous reflection?
Reflection can be computationally heavier than a normal recall operation.
Instead of making the customer-facing response wait, the system performs reflection when the ticket is resolved.
Ticket Resolution
β
βΌ
Async Reflection
β
βΌ
Hindsight Opinion
β
βΌ
Updated Risk Context
This keeps the live support interaction separate from the heavier memory-consolidation process.
2. Risk Profile Classifier
The risk engine evaluates:
- Contact frequency
- Distress indicators
- Sentiment trend
def compute_risk_profile(
self,
customer_id: str,
threads_count: int,
messages_text: str = ""
) -> RiskProfile:
"""Derives Customer Effort Trajectory
& RiskProfile with strict invariant assertions."""
contact_count = max(1, threads_count)
distress_keywords = [
"broken",
"damaged",
"again",
"still waiting",
"never arrived",
"refund",
"frustrated"
]
matches = sum(
1 for kw in distress_keywords
if kw in messages_text.lower()
)
if matches >= 1 and contact_count >= 2:
sentiment_trend = "declining"
confidence = min(
0.95,
0.78 + (contact_count * 0.04)
)
else:
sentiment_trend = "stable"
confidence = 0.80
# Evaluate risk level strictly
# according to specification thresholds
if (
contact_count >= 3
and sentiment_trend == "declining"
):
risk_level = "escalate"
elif (
contact_count >= 2
or sentiment_trend == "declining"
):
risk_level = "watch"
else:
risk_level = "normal"
assert not (
risk_level == "escalate"
and sentiment_trend != "declining"
), (
"Escalate tier requires declining sentiment"
)
return RiskProfile(
customer_id=customer_id,
contact_count_this_issue=contact_count,
sentiment_trend=sentiment_trend,
confidence=confidence,
risk_level=risk_level
)
The assertion is particularly important:
assert not (
risk_level == "escalate"
and sentiment_trend != "declining"
)
It ensures that the Escalate tier cannot be assigned without declining sentiment.
3. Multi-Session Frustration Trajectory
The system also calculates how frustration changes across individual support sessions.
def compute_frustration_trajectory(
self,
customer: Customer
) -> FrustrationTrajectory:
"""Computes multi-session frustration progression
across historical threads & live ticket."""
all_sessions = []
threads = customer.threads or []
distress_keywords = [
"broken",
"damaged",
"again",
"still waiting",
"never arrived",
"refund",
"frustrated"
]
for idx, t in enumerate(threads, 1):
text = " ".join(
[m.text for m in t.messages]
).lower()
matches = sum(
1 for kw in distress_keywords
if kw in text
)
score = min(
98,
max(
20,
25 + (idx * 18) + (matches * 12)
)
)
level = (
"Critical"
if score >= 80
else (
"High"
if score >= 60
else "Medium"
)
)
all_sessions.append(
FrustrationSession(
session_id=t.thread_id,
session_label=f"Session #{idx}",
frustration_score=score,
frustration_level=level
)
)
current_score = (
all_sessions[-1].frustration_score
if all_sessions
else 25
)
overall_trend = (
"increasing"
if (
len(all_sessions) >= 2
and (
current_score
- all_sessions[0].frustration_score
>= 15
)
)
else "stable"
)
return FrustrationTrajectory(
customer_id=customer.customer_id,
overall_trend=overall_trend,
current_frustration_score=current_score,
current_frustration_level=(
all_sessions[-1].frustration_level
if all_sessions
else "Low"
),
sessions=all_sessions
)
The resulting trajectory provides a session-by-session view:
Session #1 βββΊ Medium
β
Session #2 βββΊ High
β
Session #3 βββΊ Critical
β
βΌ
Increasing Trend
This gives the support representative a way to see how the customer's experience is changing, rather than only viewing the latest message.
4. Proactive Escalation Banner
When the risk profile reaches the escalate tier and memory is enabled, the React frontend displays an alert.
export default function EscalationBanner({
riskProfile,
memoryEnabled
}) {
if (
!memoryEnabled ||
riskProfile?.risk_level !== 'escalate'
) {
return null;
}
const contactCount =
riskProfile?.contact_count_this_issue || 3;
const confidence = Math.round(
(riskProfile?.confidence || 0.88) * 100
);
return (
<div className="bg-red-50 border-b border-red-200
border-l-4 border-l-red-600 p-4
flex items-start gap-3.5 text-red-950">
<div className="p-2 rounded-lg bg-red-100
text-red-700 mt-0.5">
<ShieldAlert className="w-5 h-5 text-red-700" />
</div>
<div>
<h4 className="text-xs font-bold
text-red-900 uppercase tracking-wider">
Proactive Escalation Alert
<span className="text-[11px]
font-semibold px-2 py-0.5 rounded-full
bg-red-100 text-red-800">
Confidence: {confidence}%
</span>
</h4>
<p className="text-xs text-red-900
mt-1.5 leading-relaxed">
Customer has contacted support
<strong>{contactCount} times</strong>
with a declining sentiment trajectory.
Prompt proactive manager intervention
or goodwill credit is strongly advised.
</p>
</div>
</div>
);
}
The banner is intentionally shown only when:
Memory = ON
+
Risk = ESCALATE
β
Proactive Escalation Alert
Memory OFF vs Memory ON
Consider Customer #4471.
The customer has already experienced two failed delivery attempts for an Echo Dot and opens a third ticket:
βWhere is my package?β
Memory Comparison
Memory OFF β Stateless Mode
Without Hindsight reflections, the LLM treats the ticket as a standard initial inquiry:
βHello Customer #4471, thanks for reaching out! Please provide your 17-digit Order ID and confirm your delivery address so I can check tracking for you.β
Result
The customer is asked to repeat information they have already provided.
Memory ON β Hindsight-Grounded Mode
With memory enabled, Hindsight reflection surfaces:
3 repeat contacts
+
Declining sentiment
β
risk_level = escalate
The copilot displays the Proactive Escalation Banner and generates a prioritized response:
βHello Customer #4471, I am very sorry to see this is your 3rd contact regarding your Echo Dot delivery (Order #302-8220-4471). I see carrier delivery failures occurred earlier this week. I have contacted carrier dispatch for priority morning redelivery and applied a $15 courtesy credit to your account.β
The architectural difference is that the second mode can use the customer's historical trajectory rather than treating the current ticket in isolation.
Usage Over Time
The dashboard provides visibility into memory operations and their usage over time, including operations such as:
- Retain
- Recall
- Reflect
- Memory/knowledge retrieval
- Memory/knowledge refresh
This provides an operational view of how the memory layer is being used by the support copilot.
What I Learned
01 β Decouple Heavy Reflection Tasks
Running areflect asynchronously during ticket resolution keeps the active message path separate from background opinion generation.
02 β Enforce Strict Risk Invariants
Sentiment analysis alone can be noisy.
Requiring:
Contact Count β₯ 3
AND
Declining Sentiment
before reaching the escalation tier provides a stricter trigger.
03 β Combine Working Memory With Temporal Memory
In-memory pinned facts can provide immediate context for specific customer constraints.
Hindsight can maintain longer-term experiences and effort trajectories.
Together:
Working Memory
+
Temporal Memory
β
Richer Customer Context
04 β Keep Support Representatives in Control
Escalation alerts and action recommendations should remain advisory.
The representative should retain explicit approval control over actions such as manager escalation or goodwill resolution.
This reduces rep cognitive load while keeping humans involved in consequential decisions.
The Complete Flow
The architecture can be summarized as:
Customer Message
β
βΌ
FastAPI Backend
β
ββββββββββββββββΊ Hindsight Recall
β β
β βΌ
β Customer History
β
βΌ
Risk & Trajectory Engine
β
βΌ
Groq LLM
β
βΌ
Agent Response
β
βΌ
Ticket Resolved
β
βΌ
Async Hindsight Reflection
β
βΌ
Updated Customer Opinion
β
βΌ
Future Escalation Signal
This creates a feedback loop:
Recall β Respond β Resolve β Reflect β Update Risk
Final Takeaway
Proactive support escalation is not simply about detecting negative words.
It requires understanding what happened across multiple interactions.
The architecture combines:
Scoped Memory
+
Temporal Experiences
+
Hindsight Reflection
+
Risk Invariants
+
Frustration Trajectory
β
Proactive Support Signals
The key idea is simple:
Don't wait for a customer to ask for a manager before recognizing that the support experience is deteriorating.
By combining persistent memory with reflection and explicit risk thresholds, the support copilot can surface relevant escalation signals earlier while keeping the final decision with the support representative.
Resources
- Hindsight β GitHub Repository
- Hindsight β Official Documentation
- Hindsight β Quickstart
- Hindsight β Guides
- Hindsight β What Agent Memory Really Means
- FastAPI β Official Documentation
- React β Official Documentation
- Groq β Official Documentation
- Groq β API Reference
- Groq β Python / JavaScript Libraries
- Vectorize β Hindsight
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.

