How We Built an AI Customer Service Agent That Cut Response Time from 45 min to 28 sec
I work on an e-commerce platform. Every day, we get about 5,000 customer service tickets. Last month, our average first response time hit 45 minutes, and our NPS was dropping faster than we could hire new agents. This i
I work on an e-commerce platform. Every day, we get about 5,000 customer service tickets. Last month, our average first response time hit 45 minutes, and our NPS was dropping faster than we could hire new agents.
This is the story of how we used Solon AI Harness to build a ticket-handling agent that cut first response time to under 30 seconds and automated 70% of our workload.
The Problem: More Tickets, Same Team
Our support team was drowning. Ticket volume was growing 15% month over month, but headcount could only grow 3%. The math didn't work.
Worse, the tickets themselves were repetitive — 70% fell into one of five categories:
- Missing delivery — customer didn't receive their package
- Wrong item — received something they didn't order
- Refund status — asking where their money is
- Return instructions — how to send something back
- General inquiry — shipping times, stock availability
Each one required an agent to open the order system, check logistics, cross-reference policies, then type a response. Every. Single. Time.
Why We Chose Harness Over a Chatbot
A simple chatbot wasn't going to cut it. These tickets need tool access — looking up orders, checking shipment status, initiating refunds. A chatbot that can't do those things is just a friendly face telling you "we'll get back to you."
Solon Harness stood out because it's built specifically for agents that need to do things, not just say things. The architecture is simple:
Harness (runtime) → Agent (decision maker) → Tools (actions)
↑
Context/Memory (state)
Each ticket becomes a conversation session. The agent decides which tool to call based on what it sees, and Harness handles the loop — call tool, get result, decide what's next, repeat until resolved.
How We Built It
Step 1: Define the Tools
The agent needed three tools to do its job:
@Slf4j
public class OrderTool implements Tool {
@Override
public String invoke(String args) throws Exception {
// Look up order by ID → return status, items, price
Order order = orderService.findById(args.trim());
return JsonUtils.toJson(order);
}
}
@Slf4j
public class LogisticTool implements Tool {
@Override
public String invoke(String args) throws Exception {
// Check logistics tracking → return current status
LogisticStatus status = logisticService.track(args.trim());
return JsonUtils.toJson(status);
}
}
@Slf4j
public class CompensationTool implements Tool {
@Override
public ToolMetadata getMetadata() {
return ToolMetadata.of("apply_compensation",
"Apply refund or resend for a qualified order");
}
@Override
public String invoke(String args) throws Exception {
CompensationRequest req = JsonUtils.toObject(args, CompensationRequest.class);
return compensationService.apply(req);
}
}
Three tools, about 60 lines each. The hardest part wasn't writing them — it was deciding which ones to build first.
Step 2: Connect Harness
Harness harness = Harness.harness(new ReActAgent(chatModel))
.addTools(
new OrderTool(orderService),
new LogisticTool(logisticService),
new CompensationTool(compensationService)
)
.addInterceptor(new StopLoopInterceptor(10)) // max 10 turns
.start();
That's it. The interceptor prevents runaway loops — if the agent hasn't resolved the ticket in 10 tool calls, it escalates to a human.
Step 3: The "Safe Refund" Pattern
This was the part that made our compliance team nervous. What if the agent refunds an order that shouldn't be refunded?
Harness has a Human-in-the-Loop (HITL) pattern built in. We wrapped high-risk operations:
// CompensationTool with HITL guard
if (req.amount() > 100) {
// Escalate for manual approval
escalationService.create(req, currentSession);
return "⏳ Refund >$100 requires manual approval. Escalated to team lead.";
}
// Auto-process small refunds
return compensationService.apply(req);
Small refunds (< $100) go through automatically. Larger ones get flagged with full context — who the customer is, what happened, what the agent recommends. The human just clicks approve or deny.
The Results After 4 Weeks
We rolled this out incrementally, starting with just "missing delivery" tickets, then expanding to all five categories.
| Metric | Before | After | Improvement |
|---|---|---|---|
| First response time | 45 min | 28 sec | 96x faster |
| Auto-resolution rate | 0% | 72% | New capability |
| Human agent workload | 5,000 tickets/day | 1,400 tickets/day | 72% reduction |
| Customer satisfaction | 3.2/5 | 4.1/5 | +0.9 pts |
| Escalation accuracy | N/A | 94% | 6% false positives sent to humans |
The 6% that should have been auto-resolved but got escalated? We tuned the prompts and added one more tool (a "policy lookup" tool that checks return windows). By week 3, false escalations dropped to under 2%.
What Surprised Us
Three things we didn't expect:
1. Agents handled ambiguity better than we thought. A ticket that said "I didn't get my stuff" — the agent checked the order, found it was delivered 3 days ago, then asked the customer to check with neighbors. A human would have done the exact same thing.
2. The tool permission granularity mattered more than we guessed. We initially gave all tools to all agents. Bad idea. We ended up with three agent tiers — Tier 1 (read-only: order + logistics queries), Tier 2 (auto-refund up to $100), Tier 3 (full access + HITL). Each tier is just a different tool configuration.
3. Session context is everything. Customers get frustrated when they have to repeat themselves. Harness's session memory means the agent remembers the entire conversation across tool calls. If it checked the order in step 2, it doesn't re-ask in step 5.
What I'd Do Differently
Looking back, I'd start with even fewer tools. We deployed with six tools initially, and two of them (inventory lookup and coupon application) were almost never used. The agent was making fine-grained decisions with tools it didn't need, adding latency and token cost. Three tools did the job.
Also — invest in the escalation UI early. The tool calls work great, but when a ticket escalates to a human, they need to see the agent's reasoning trail. We built this in week 3 and it was a game-changer for agent trust.
Why This Matters
Customer service is where AI agents deliver the most obvious ROI today. It's not glamorous — no one writes blog posts about refund processing — but it's the kind of work that pays for itself in weeks, not months.
The Solon Harness made this possible because it gave us the exact capabilities we needed and nothing we didn't:
- Tool system — the agent could actually do things
- Loop control — no runaway costs
- HITL — compliance slept well at night
- Session memory — no repeat questions
- Lightweight runtime — 0.3m kernel, started in under a second
If you're dealing with high-volume customer tickets and thinking about AI, skip the chatbot. Build an agent that can actually resolve things. It's not as hard as you think.
Built with Solon AI Harness 4.0.3. The full source code for this project is available on GitHub.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.