Dev.to WebDev ๐Ÿ›  Dev ๐Ÿ‘ 0 ๐Ÿ“– 13 min read

How to Debug "Laggy" Live Scores: Finding Where Your Latency Is Actually Coming From

Your users say the live score is "laggy". Your first instinct is probably to blame the API. Your API provider's first instinct is to blame your code. Both of you may be wrong. "Laggy" is not a measurement. It's a feelin

Your users say the live score is "laggy". Your first instinct is probably to blame the API. Your API provider's first instinct is to blame your code. Both of you may be wrong.

"Laggy" is not a measurement. It's a feeling, and that feeling can come from at least eight different places: DNS, TLS handshakes, a polling interval, a CDN cache, a blocked main thread, a backgrounded browser tab, a slow webhook queue, or the broadcast itself being 20 seconds behind the stadium (more on that last one later).

This post is a practical debugging playbook. We'll instrument every hop between "something happened on the pitch" and "a pixel changed on screen", then use the numbers to find the real culprit. Code examples use the Orbistats Live Scores API, but the technique works with any provider.

Note on numbers: I don't quote benchmark results in this post. Every table is a template you fill with your own measurements.

Table of contents
First, define "laggy" precisely
Map the pipeline
Step 1: Rule out the network layer with curl
Step 2: Check your polling interval math
Step 3: Add per-hop timestamps
Step 4: Fix your clocks before trusting any number
Step 5: Measure delivery over WebSocket
Step 6: Catch main-thread and render lag in the browser
Step 7: Webhook pipelines and queue delay
The "not actually lag" cases
A diagnostic decision tree
Build a staleness monitor
Final checklist and next steps

  1. First, define "laggy" precisely

Before touching code, get a reproducible complaint. Ask for or capture:

Which match, which event (goal, card, score change)?
What time did they see it, and what did a reference source show?
Which device, browser and network (Wi-Fi, mobile data)?
Is the lag constant, or does it spike?

"Constant 10 seconds late" and "occasionally 3 seconds late" are different bugs. Constant offset usually means a polling interval or caching layer. Spikes usually mean reconnects, garbage collection, a throttled tab or network jitter.

  1. Map the pipeline

Every live score travels through the same stages:

text
Real-world event
โ†“ (A) data source / provider ingestion
Provider system
โ†“ (B) provider delivery: REST / WebSocket / Webhook
Network
โ†“ (C) your ingest server
Your backend (parse, cache, fan-out)
โ†“ (D) your delivery to clients
Network again
โ†“ (E) browser/app receives bytes
Client JS (parse, state update)
โ†“ (F) render
Pixel on screen

Providers control A and B. You control C through F. Your goal is to put a timestamp at every arrow so you can subtract and see which segment is fat. Orbistats positions its live feed as a sub-50ms system, and it has written about what sub-50ms actually requires, end to end. That claim concerns the provider side. Everything after it is yours.

Here's the table to fill in:

Segment What it covers Your measured p50 Your measured p95
Aโ†’B Event โ†’ provider emits ? ?
Bโ†’C Provider โ†’ your server ? ?
Cโ†’D Your server โ†’ your fan-out ? ?
Dโ†’E Your server โ†’ client bytes ? ?
Eโ†’F Client bytes โ†’ pixel ? ?

  1. Step 1: Rule out the network layer with curl

Start with the cheapest test. curl can break a single request into its phases:

bash
cat > curl-format.txt <<'EOF'
dns: %{time_namelookup}s
tcp: %{time_connect}s
tls: %{time_appconnect}s
ttfb: %{time_starttransfer}s
total: %{time_total}s
size: %{size_download} bytes
EOF

curl -s -o /dev/null -w "@curl-format.txt" \
-H "Authorization: Bearer $ORBISTATS_KEY" \
https://api.orbistats.com/v1/football/matches/live

How to read it (the values are cumulative, so subtract neighbours):

dns high โ†’ resolver problem; try a different DNS or cache lookups.
tcp โˆ’ dns โ†’ raw network distance to the server.
tls โˆ’ tcp โ†’ handshake cost. If this appears on every request, you're not reusing connections.
ttfb โˆ’ tls โ†’ server think time plus the first-byte trip.
total โˆ’ ttfb โ†’ payload download time; large for big responses.

Run it 20 times, not once:

bash
for i in $(seq 1 20); do
curl -s -o /dev/null -w "%{time_starttransfer}\n" \
-H "Authorization: Bearer $ORBISTATS_KEY" \
https://api.orbistats.com/v1/football/matches/live
done | sort -n

The sorted list gives you a feel for median and tail. If the first request is slow and the rest are fast, that's connection setup, which you can fix with keep-alive.

Fix: reuse connections

In Node 20+, the built-in fetch (undici) already pools connections, but a naive script that spawns a new process per poll never benefits. If you're on an older HTTP client, enable keep-alive explicitly:

js
import https from "node:https";
const agent = new https.Agent({ keepAlive: true, maxSockets: 10 });
// pass { agent } to your requests

Not sure about the exact live endpoint for a given sport? The API reference lists every route, and the Sandbox lets you try requests in the browser first.

  1. Step 2: Check your polling interval math

If you poll, the single biggest source of "lag" is simply the interval. It's arithmetic, not a bug:

Average added delay = interval / 2
Worst case = interval
Poll interval Average staleness Worst case
30 s 15 s 30 s
10 s 5 s 10 s
5 s 2.5 s 5 s
1 s 0.5 s 1 s

Now stack the layers. This is the classic trap:

text
Your frontend polls your backend every 10 s
Your backend polls the provider every 10 s
โ†“
Worst case staleness = 10 s + 10 s = 20 s
Average = 5 s + 5 s = 10 s

Two chained pollers add their staleness. If users report "about 10 seconds behind", check for exactly this setup before suspecting anything exotic.

Also check request budgets. A plan with a daily request cap (see the pricing page for current limits) can't sustain a 1-second poll for long. If you silently hit a 429 and your code keeps showing the last cached score, the UI looks frozen. Log every non-200:

js
const res = await fetch(url, { headers });
if (!res.ok) {
console.warn("poll failed", res.status, res.headers.get("retry-after"));
}

  1. Step 3: Add per-hop timestamps

Now we instrument. The idea is to stamp the message at every hop and carry the stamps forward, so the browser can compute every segment.

js
// ingest.js - runs on YOUR server when data arrives from the provider
function stampIngest(providerMsg) {
return {
...providerMsg,
_t: {
// provider's own event time, if present in the payload (check the docs for the field name)
provider: providerMsg.timestamp ?? null,
ingest: Date.now(), // when YOUR server received it
},
};
}

// gateway.js - just before sending to browsers
function stampEmit(msg) {
msg._t.emit = Date.now(); // when YOUR server sent it to clients
return msg;
}

And in the browser:

js
// client.js
socket.onmessage = (e) => {
const msg = JSON.parse(e.data);
msg._t.recv = Date.now(); // bytes arrived

requestAnimationFrame(() => {
paint(msg);
msg._t.paint = Date.now(); // pixel updated (approx.)
report(msg._t);
});
};

function report(t) {
const seg = {
provider_to_ingest: t.provider ? t.ingest - t.provider : null,
ingest_to_emit: t.emit - t.ingest,
emit_to_recv: t.recv - t.emit, // network + client clock skew!
recv_to_paint: t.paint - t.recv,
total: t.paint - (t.provider ?? t.ingest),
};
navigator.sendBeacon("/metrics/latency", JSON.stringify(seg));
}

Ship these segments to whatever metrics store you use (even a simple log line works) and aggregate p50 / p95 / p99 per segment. The fattest segment is where you look next.

Don't average. If p50 is fine and p99 is terrible, users remember the p99.

  1. Step 4: Fix your clocks before trusting any number

This is the step that makes or breaks the whole exercise. emit_to_recv compares a server clock with a browser clock. If those disagree by 400 ms, your "network latency" is off by 400 ms, and could even go negative.

Estimate the client's offset against your server with an NTP-style handshake. Use your own endpoint that returns milliseconds (the HTTP Date header only has one-second resolution, which is too coarse):

js
// server: GET /time -> { now: Date.now() }
app.get("/time", (_req, res) => res.json({ now: Date.now() }));
js
// client
async function estimateOffset(samples = 8) {
const results = [];
for (let i = 0; i < samples; i++) {
const t0 = Date.now();
const { now: serverNow } = await (await fetch("/time", { cache: "no-store" })).json();
const t1 = Date.now();
const rtt = t1 - t0;
const offset = serverNow - (t0 + rtt / 2); // assumes symmetric path
results.push({ rtt, offset });
}
// Trust the sample with the smallest RTT: least queueing noise.
results.sort((a, b) => a.rtt - b.rtt);
return results[0].offset;
}

const offset = await estimateOffset();
// later: serverTimeNow โ‰ˆ Date.now() + offset

Then correct your measurements: emit_to_recv = (recv + offset) - emit.

On your servers, make sure NTP/chrony is running. Two machines with drifting clocks will produce phantom latency between your ingest and gateway nodes too.

  1. Step 5: Measure delivery over WebSocket

If you already stream, "lag" usually has a different set of suspects than polling. Check these in order:

(a) Is the socket actually alive? A half-open TCP connection looks connected but delivers nothing, for minutes. Use heartbeats and a "last message seen" watchdog:

js
let lastMsgAt = Date.now();

ws.on("message", () => { lastMsgAt = Date.now(); });

setInterval(() => {
if (Date.now() - lastMsgAt > 30_000) {
console.warn("socket silent for 30s, forcing reconnect");
ws.terminate(); // triggers your reconnect logic
}
}, 5_000);

During live play on a busy match, 30 seconds of silence is already suspicious. Tune the threshold to your sport: a golf round has long quiet stretches, while a basketball game rarely does.

(b) Are you applying backpressure correctly? If your handler does slow work (database writes, heavy JSON) inside the message callback, messages queue up behind it and every later update looks "late". Keep the handler tiny:

js
ws.on("message", (raw) => {
const msg = JSON.parse(raw);
latestState.set(msg.match_id, msg); // O(1) state write
scheduleBroadcast(); // defer everything else
});

(c) Did you miss updates after a reconnect? After any reconnect, fetch a REST snapshot, then resume the stream. Otherwise the score stays wrong until the next change arrives. Connection and subscribe details are on the WebSocket API page.

(d) Are you flooding the client? Ten updates in one frame should paint once, not ten times:

js
const pending = new Map();
let raf = 0;

function enqueue(msg) {
pending.set(msg.match_id, msg); // keep only the latest per match
if (!raf) {
raf = requestAnimationFrame(() => {
raf = 0;
for (const m of pending.values()) paint(m);
pending.clear();
});
}
}

  1. Step 6: Catch main-thread and render lag in the browser

Sometimes the data arrives instantly and the UI still feels slow. That's a client problem. The browser can tell you directly.

Find long tasks (anything blocking the main thread for 50 ms or more):

js
new PerformanceObserver((list) => {
for (const e of list.getEntries()) {
console.warn(long task: ${e.duration.toFixed(0)}ms, e);
}
}).observe({ entryTypes: ["longtask"] });

If a long task coincides with each score update, your render path is the problem. Common culprits: re-rendering a whole match list instead of one row, huge unvirtualized lists, synchronous JSON parsing of large payloads, and expensive CSS animations.

Use the Performance panel, not guesses. Record while a live update lands. The flame chart will show whether the time is spent in scripting, layout or paint.

Handle hidden tabs. This one fools a lot of people. Browsers throttle background tabs: timers get delayed, requestAnimationFrame stops firing entirely, and on some browsers the throttling becomes aggressive after a few minutes. A user who switches tabs and comes back sees stale data for a moment. Fix it by listening for visibility and resyncing:

js
document.addEventListener("visibilitychange", async () => {
if (document.visibilityState === "visible") {
const snapshot = await fetch("/api/live-snapshot").then((r) => r.json());
applySnapshot(snapshot); // jump straight to current truth
}
});

  1. Step 7: Webhook pipelines and queue delay

Webhooks give you push-style delivery to your server, but they have their own lag sources, and none of them are visible from the provider's side.

Slow acknowledgements. If your endpoint does heavy work before replying 200, the provider may time out and retry, which creates duplicates and delay.
Queue buildup. If you push events onto a queue and your consumers fall behind during a busy match window (think a full Saturday of football fixtures), queue depth becomes latency. Monitor queue age, not just queue size.
Cold starts. A serverless function that sleeps between goals adds a startup penalty exactly when the goal arrives.
Retries. Out-of-order delivery after a retry can make an old score overwrite a new one.

Guard against the last one by comparing event time or a sequence value before applying an update:

js
function applyIfNewer(state, msg) {
const current = state.get(msg.match_id);
// 'seq' or an event timestamp: use whichever the payload actually provides
if (current && current.seq >= msg.seq) return; // stale, ignore
state.set(msg.match_id, msg);
}

Signature checks, retry behaviour and payload shapes are documented on the Webhooks API page. Always acknowledge first and process after.

  1. The "not actually lag" cases

Before you spend a week optimizing, rule out these:

The broadcast is behind. TV and streaming broadcasts are often many seconds behind the real event, and different services are behind by different amounts. If a user compares your app to a live TV stream, a data feed that is ahead can look like it's "spoiling" rather than lagging, and a feed compared against a faster social post can look slow. Always compare against a consistent reference, ideally the same data source on a stopwatch.

Different definitions of "final". One site updates the score on the goal, another waits for VAR confirmation. Two correct systems can disagree for a minute.

Cache headers. A CDN or browser cache serving a 30-second-old response looks exactly like lag. Check the response headers:

bash
curl -sI https://your-domain.com/api/live-snapshot | grep -iE "cache-control|age|etag|cf-cache-status|x-cache"

If age is large or cache-control: max-age=30 is set on a live endpoint, you've found it. For live data, use no-store, or a very short s-maxage deliberately.

Region distance. A server in one continent serving users in another adds real round-trip time to every non-streamed request. Measure from where your users actually are.

  1. A diagnostic decision tree

When someone reports lag, walk this tree top to bottom and stop at the first match:

text
Is the score wrong for a long time after a change (many seconds)?
โ”œโ”€ YES โ†’ Is there a cache header or CDN in front? โ†’ fix caching
โ”‚ Is there a poller (frontend or backend)? โ†’ check interval math (ยง4)
โ”‚ Did the socket go silent / reconnect? โ†’ watchdog + snapshot (ยง7)
โ”‚ Did a request return 429/5xx? โ†’ rate limits / status page
โ””โ”€ NO โ†’ Does it update fast but "feel" janky?
โ”œโ”€ Long tasks in the browser? โ†’ optimize render (ยง8)
โ”œโ”€ Only after switching tabs? โ†’ visibilitychange resync (ยง8)
โ””โ”€ Compared against TV/other site? โ†’ reference mismatch (ยง10)

When you suspect the provider rather than your own stack, check the Status page and the changelog first, so you can separate "an incident is ongoing" from "my code regressed".

  1. Build a staleness monitor

The best defense is catching lag before users do. A staleness monitor tracks "how old is the freshest data for each live match" and alerts when it exceeds a threshold:

js
// monitor.js
const lastUpdate = new Map(); // match_id -> ms timestamp of last data
const LIVE = new Set(); // match_ids currently in play

export function onData(matchId) {
lastUpdate.set(matchId, Date.now());
}

setInterval(() => {
const now = Date.now();
for (const id of LIVE) {
const age = now - (lastUpdate.get(id) ?? 0);
if (age > 20_000) {
alert(match ${id}: no update for ${Math.round(age / 1000)}s);
}
}
}, 5_000);

function alert(msg) {
console.error("[STALENESS]", msg);
// send to Slack / PagerDuty / your logger here
}

Pair it with the per-segment metrics from section 5 and a simple dashboard with three lines per segment (p50, p95, p99). Within a week you'll know your real latency profile instead of arguing about feelings.

  1. Final checklist and next steps Reproduce the complaint with a specific match, event and device curl -w breakdown: DNS, TCP, TLS, TTFB, total Reuse connections (keep-alive) Do the interval math, and check for chained pollers Stamp every hop: provider, ingest, emit, receive, paint Correct clock offset on client and sync servers with NTP Add a "socket silent" watchdog and a REST snapshot after reconnect Coalesce updates with requestAnimationFrame Observe long tasks; resync on visibilitychange Verify cache headers on every live endpoint Alert on staleness, not only on errors Log non-200 responses, especially 429

If you want a clean environment to practice on:

Create a free key on the signup page.
Fire your first request from the Quickstart, then browse the documentation for the full resource list.
Test routes in the Sandbox before wiring them into code.
Skipping custom UI altogether? The embeddable widgets handle rendering for you.
Compare plans and request limits on the pricing page.

Coverage spans 13 sports, so the same debugging approach applies whether you're tracking football or cricket.

What was the weirdest source of "lag" you've ever tracked down? A cached CDN response, a throttled tab, a chained poller? Share it in the comments. I'm collecting war stories.

๐Ÿ“ฐ Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.