5 Mistakes That Make LLM Streaming Break on a Flaky Network
Your chat UI looks fine on the office Wi-Fi. Then someone joins from a train, the SSE connection drops mid-sentence, and either the answer restarts from scratch or the model keeps burning tokens while nobody is listening
Your chat UI looks fine on the office Wi-Fi. Then someone joins from a train, the SSE connection drops mid-sentence, and either the answer restarts from scratch or the model keeps burning tokens while nobody is listening. Same feature, different network.
I spent a stretch of time turning a "works in demos" stream into something that survives reconnects. The full working project (ASP.NET Core 10, TypedResults.ServerSentEvents, tests) lives in Tech Skill Builder. This post is the short checklist: five mistakes I kept seeing, and the shape that fixes them.
Mistake 1: Tying generation to the HTTP request
The five-minute version wires GetStreamingResponseAsync to HttpContext.RequestAborted. Close the tab, and the model call stops. That feels thrifty until you realize a flaky proxy is doing the same thing: every blink cancels work you already paid for, and a reconnect starts a new prompt.
Fix: split start from subscribe.
-
POST /api/chat/streamsstarts generation in a registry and returns202withstreamId,eventsUrl,cancelUrl. -
GET /api/chat/streams/{id}/eventsis the SSE subscription. - Generation keeps running even if the subscriber disconnects (until a sweeper decides nobody cares anymore).
Skeleton:
// POST /api/chat/streams
var stream = registry.TryStart(messages, owner);
return TypedResults.Accepted(
$"/api/chat/streams/{stream.Id}/events",
new StreamStarted(stream.Id, eventsUrl, cancelUrl));
// GET /api/chat/streams/{id}/events
// Last-Event-ID header (or ?lastEventId=) รขโ โ replay missed events
return TypedResults.ServerSentEvents(
StreamSubscriber.ReadAsync(stream, after, options, time, http.RequestAborted));
// full implementation in the complete project
EventSource only speaks GET. That alone pushes you toward this split: you cannot POST a body on reconnect.
Mistake 2: No sequence ids, so reconnect means "start over"
Without an SSE id: on each event, the browser has nothing to put in Last-Event-ID. Your client either duplicates tokens or asks the server to regenerate.
Fix: keep an append-only StreamLog per answer. Every event gets a sequence number. On subscribe, parse Last-Event-ID, then replay only what was missed. Bound the buffer; a very late client gets one snapshot with the text so far instead of every erased delta.
public long Append(ChatStreamEvent e)
{
// assign next sequence, store event, wake waiters
// if over capacity รขโ โ trim oldest; keep aggregated text for snapshots
// full implementation in the complete project
}
Caught-up client on a finished stream? Return 204 No Content. That is the SSE-spec way to stop EventSource from reconnecting forever.
Mistake 3: Naming a server event error
Browsers fire EventSource's own error for connection trouble. If your payload also uses event: error, both land in the same mental bucket (and often the same handler). You will debug "is the model broken or is the Wi-Fi broken?" for longer than you want.
Fix: use a small, boring vocabulary:
| Event | Meaning |
|---|---|
delta |
text fragment |
tool |
name + status only (not arguments/results) |
done |
finished (stop, cancelled, รขโฌยฆ) |
failed |
in-band model/server failure after HTTP 200 |
heartbeat |
keep proxies from killing an idle stream |
snapshot |
replace local text after a truncated buffer |
Call the failure event failed, not error. Put a user-safe message and an errorId in the payload. Keep partial text on screen.
Mistake 4: Silent streams and no Stop button
Proxies and load balancers love idle connections. While the model thinks or a tool runs, your SSE can sit quiet long enough to get cut. Separately, users hit Stop and expect the bill to stop too. If cancel only closes the browser side, the server may keep generating.
Fix:
- Emit
heartbeaton a timer (for example every 15 seconds) while work is in progress. - Expose
POST /api/chat/streams/{id}/cancelthat cancels the model call and ends withdone/finishReason: cancelled. - Send
X-Accel-Buffering: noso nginx does not buffer the stream into one giant blob.
// Cancel endpoint shape
if (registry.Find(id) is not { } stream || stream.Owner != owner)
return TypedResults.NotFound();
return stream.RequestCancel()
? TypedResults.Accepted((string?)null)
: TypedResults.Conflict(); // already finished
// full implementation in the complete project
Mistake 5: Shipping the "quick" stream as production
TypedResults.ServerSentEvents over a raw token loop is great for a spike. It is not a product API. You get no resume, no ownership check, no rate limit on active streams, and no story for multi-instance hosts.
Checklist before you call it done:
- [ ] POST start + GET subscribe (EventSource-friendly)
- [ ] Sequence ids +
Last-Event-IDreplay - [ ] Generation outlives the request; sweeper cleans orphans
- [ ] Event names avoid colliding with EventSource (
failednoterror) - [ ] Heartbeats + cancel endpoint
- [ ] Owner binding (cookie auth or signed URL; EventSource cannot set
Authorization) - [ ] Sticky sessions or a shared log (Redis Streams, etc.) if you scale out
- [ ] Prefer HTTP/2 or HTTP/3 (browsers cap ~6 SSE connections per domain on HTTP/1.1)
What I leave out here on purpose
The complete project has the coalescer (batch tiny token fragments), the sweeper timings, the offline streaming model for demos without a key, and the full test suite. This article is the map, not the hike.
If you want the working source and the deeper write-up package, grab it from Tech Skill Builder. Membership includes the full resumable streaming solution ready to run with dotnet test and dotnet run. Limited-time member pricing is on the product page; if you are building anything that streams tokens to a browser, this is the piece you want finished before the next flaky demo.
Flaky networks are not an edge case. Treat resume as part of the feature, not a patch after the first support ticket.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.