Dev.to AI 🤖 Ai 👁 0 📖 5 min read

Implementing AI Observability: Tracing LLM Calls End-to-End

Your LLM-powered app is in production. Users are hitting it. Something is slow — or wrong — and you have no idea what. You can't reproduce it locally, and your generic APM shows... a single HTTP call to an external API.

Your LLM-powered app is in production. Users are hitting it. Something is slow — or wrong — and you have no idea what. You can't reproduce it locally, and your generic APM shows... a single HTTP call to an external API. That's the problem with LLM observability today: the gap between "a request happened" and "here's what the model did, at what cost, with what latency" is usually empty.

This article shows how to close that gap with a tracing layer you can build in an afternoon.

What You Need to Trace

Before writing any code, define what actually matters for LLM calls:

  • Latency breakdown: time to first token (TTFT) vs. total completion time
  • Token usage: prompt tokens, completion tokens, cached tokens (if your provider supports it)
  • Cost: calculated per call, not just per month
  • Model and version: which model handled the request
  • Input/output sampling: truncated copies of the prompt and response for debugging
  • Errors: rate limits, timeouts, content policy rejections

This is different from what a generic APM captures. OpenTelemetry traces give you spans and HTTP status codes. They won't tell you "prompt tokens cost $0.0023 and the model refused because of the system prompt." You need a layer that understands LLM semantics.

Building a Wrapper Tracer in Python

The cleanest approach is a thin wrapper around your LLM client. You keep all existing call sites, add one decorator, and tracing is transparent.

import time
import json
import uuid
import logging
from dataclasses import dataclass, field
from typing import Any

logger = logging.getLogger("llm.tracer")

@dataclass
class LLMTrace:
    trace_id: str = field(default_factory=lambda: str(uuid.uuid4())[:8])
    model: str = ""
    prompt_tokens: int = 0
    completion_tokens: int = 0
    cached_tokens: int = 0
    latency_ms: float = 0.0
    ttft_ms: float = 0.0
    cost_usd: float = 0.0
    error: str | None = None
    response_preview: str = ""

# Cost per 1M tokens (input / output)
COST_TABLE = {
    "claude-opus-5":      (15.0, 75.0),
    "claude-sonnet-5":    (3.0, 15.0),
    "claude-haiku-4-5":   (0.8, 4.0),
    "gpt-4o":             (2.5, 10.0),
    "gpt-4o-mini":        (0.15, 0.6),
}

def calculate_cost(model: str, prompt_tokens: int, completion_tokens: int) -> float:
    costs = COST_TABLE.get(model, (0.0, 0.0))
    return (prompt_tokens / 1_000_000 * costs[0]) + (completion_tokens / 1_000_000 * costs[1])


def trace_llm_call(func):
    import functools

    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        trace = LLMTrace()
        start = time.monotonic()
        try:
            result = func(*args, **kwargs)
            trace.latency_ms = (time.monotonic() - start) * 1000
            _extract_from_response(trace, result)
            return result
        except Exception as e:
            trace.latency_ms = (time.monotonic() - start) * 1000
            trace.error = type(e).__name__ + ": " + str(e)[:200]
            raise
        finally:
            _emit(trace)

    return wrapper


def _extract_from_response(trace: LLMTrace, response: Any) -> None:
    usage = getattr(response, "usage", None)
    if usage:
        trace.prompt_tokens = (
            getattr(usage, "input_tokens", 0) or
            getattr(usage, "prompt_tokens", 0)
        )
        trace.completion_tokens = (
            getattr(usage, "output_tokens", 0) or
            getattr(usage, "completion_tokens", 0)
        )

    model = getattr(response, "model", "")
    trace.model = model
    trace.cost_usd = calculate_cost(model, trace.prompt_tokens, trace.completion_tokens)

    content = getattr(response, "content", None) or getattr(response, "choices", None)
    if content and isinstance(content, list):
        if hasattr(content[0], "text"):
            trace.response_preview = content[0].text[:300]
        elif hasattr(content[0], "message"):
            trace.response_preview = content[0].message.content[:300]


def _emit(trace: LLMTrace) -> None:
    logger.info(json.dumps({
        "trace_id": trace.trace_id,
        "model": trace.model,
        "prompt_tokens": trace.prompt_tokens,
        "completion_tokens": trace.completion_tokens,
        "latency_ms": round(trace.latency_ms, 1),
        "cost_usd": round(trace.cost_usd, 6),
        "error": trace.error,
    }))

You use it like this:

@trace_llm_call
def classify_text(text: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-5",
        max_tokens=256,
        messages=[{"role": "user", "content": f"Classify as spam or not: {text}"}],
    )
    return response.content[0].text

Every call now emits a structured JSON log with token counts, cost, and latency. No external SDK required.

Adding Time-to-First-Token for Streaming

For streaming responses, total latency is misleading. A call that takes 3s total but starts streaming after 200ms feels fast. One that stalls for 2.8s before the first token feels broken, even if total time is identical. Measure both separately.

def trace_streaming_call(client, model: str, messages: list[dict]) -> str:
    trace = LLMTrace(model=model)
    start = time.monotonic()
    first_token_time: float | None = None
    chunks: list[str] = []

    with client.messages.stream(
        model=model,
        max_tokens=512,
        messages=messages,
    ) as stream:
        for event in stream:
            if (first_token_time is None
                    and hasattr(event, "type")
                    and event.type == "content_block_delta"):
                first_token_time = time.monotonic()
                trace.ttft_ms = (first_token_time - start) * 1000

            if hasattr(event, "delta") and hasattr(event.delta, "text"):
                chunks.append(event.delta.text)

        final = stream.get_final_message()

    trace.latency_ms = (time.monotonic() - start) * 1000
    _extract_from_response(trace, final)
    trace.response_preview = "".join(chunks)[:300]
    _emit(trace)
    return "".join(chunks)

With TTFT tracked independently, you can alert on provider-side slowdowns separate from model generation time — a much more actionable signal.

Shipping Traces to a Backend

Structured JSON logs are a solid foundation, but for dashboards and alerting you want something queryable. Two practical paths:

OpenTelemetry (OTLP): convert each trace to an OTEL span and ship to any compatible collector (Jaeger, Grafana Tempo, Honeycomb). Vendor-neutral, but OTEL's data model wasn't designed for LLM-specific fields — cost_usd and cached_tokens end up as plain span attributes without native UI treatment.

Dedicated LLM observability tools (Langfuse, Phoenix, Helicone): they understand LLM semantics natively — token cost timelines, model comparison, prompt diffing. Most are open-source and self-hostable. If your prompts contain sensitive data, self-hosting is the only responsible option.

The pragmatic path for most teams: structured logs first (zero dependencies, works today), then add a dedicated platform once you know which metrics you actually query. Getting any visibility now beats waiting for perfect infrastructure.

For teams building secure AI systems, pairing observability with input validation and output controls is worth structuring from the start. The security hardening checklists at ayinedjimi-consultants.fr/checklists include an LLM section covering rate limiting, prompt injection detection, and output schema validation — all things that interlock naturally with the tracing layer above.

What to Alert On

Once traces are flowing, these thresholds are worth configuring first:

Metric Alert condition
Error rate > 2% over 5 minutes
P95 total latency > 5s
TTFT P50 > 1.5s
Hourly cost > budget threshold
Cached token ratio < 30%

The cost alert is the most overlooked. A loop bug in an agent pipeline can burn hundreds of dollars in minutes. An hourly budget cap with a hard circuit breaker is not optional in production.

The Takeaway

LLM observability requires instrumenting at the right semantic level — not just "an HTTP call happened" but "which model, how many tokens, was it cached, what did it cost, and how long to first token." A thin wrapper around your LLM client gives you all of that in under 100 lines of Python, with no external dependencies. Add streaming TTFT measurement early — it surfaces provider-side issues that total latency hides. Ship structured logs first, then evaluate dedicated platforms once you know which questions you actually need to answer.

I run AYI NEDJIMI Consultants, a cybersecurity consulting firm. We publish free security hardening checklists — PDF and Excel.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.