Dev.to AI 🤖 Ai 👁 0 📖 10 min read

Frontend Plus Backend Error Tracking: JavaScript and API Trace Correlation

The important trade-off is fidelity versus operational weight: use a backend error pipeline as the system of record for API failures and AI-call cost, then forward compact browser error summaries through that backend wit

The important trade-off is fidelity versus operational weight: use a backend error pipeline as the system of record for API failures and AI-call cost, then forward compact browser error summaries through that backend with the same trace_id or request_id. Short answer: this is the simplest defensible setup for a logistics agent loop when the question is "which failed shipment-planning run consumed this model cost?" It is not a substitute for source-map decoding, session replay, or a distributed trace viewer; add a browser-focused product when those are part of the debugging objective.

That boundary matters. A browser exception saying "plan generation failed" is nearly useless unless an operator can connect it to the API request, the model invocation, the carrier lookup, and the cost metadata recorded for that run. The correlation identifier provides the join key. It does not magically provide a span tree.

One run, one join.

The evidence packet for one dispatch attempt

Consider a bounded failure during a dispatch window: the agent proposes no route, the browser reports an error, the API returns a failure, and several model calls have already accrued cost. My first move in the review would be to resist counting four alerts as four incidents. I would ask for one correlation value carried from the edge request into every backend log and into the sanitized browser report, then group the evidence by agent run.

The invariant is plain: one user-visible attempt needs one stable join key, while every internal operation may have its own span identifier. A trace_id follows the full attempt; a span_id distinguishes work such as model inference or a carrier API call; a request_id can remain the narrower HTTP identifier if that is already the platform convention. Pick the semantics once. Mixing those names for the same value produces correlations that look convincing and are wrong.

For an AI loop, record cost at the call boundary rather than estimating it later from an error count. Useful event fields include the stable trace value, agent run ID, model and vendor, cost_usd, latency_ms, outcome, and a low-cardinality stage such as classify, plan, or validate. Do not put prompts, access tokens, customer addresses, or arbitrary exception text into metric labels. Logs can carry carefully redacted detail; metrics should support capacity planning without creating an unbounded series count.

One short incident can otherwise inflate three numbers independently: browser errors, API exceptions, and failed model calls. The join key lets support reconstruct the sequence, while the agent run ID lets finance attribute spend to the workflow. Keep both.

How should a React frontend plus Node.js backend correlate errors?

The backend should issue or accept a valid correlation value, return it in a response header, and include it in structured logs. The browser's global error handler can read the value retained from the relevant API response and POST a sanitized summary to an application-owned endpoint. That endpoint, not the browser, sends data onward to the chosen error service, so no ingestion credential is exposed in client code.

This runnable Go service demonstrates the server-side contract. It uses only the standard library, applies a size limit, rejects malformed input, and logs JSON that can be joined on trace_id. It also exposes an internal handler that performs a real authenticated log search against a plain REST API, with an explicit method, status checks, and bounded rate-limit retries. No search filters appear because that route's discovery parameters are undeclared; inventing a convenient trace_id query parameter would make the sample look better while teaching an unsupported contract. The browser capture code is intentionally omitted because this article's code convention is Go-only; the wire contract is the part that must remain stable across frontend frameworks.

package main

import (
    "crypto/rand"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type browserError struct {
    TraceID string `json:"trace_id"`
    Message string `json:"message"`
    Page    string `json:"page"`
}

func newTraceID() (string, error) {
    b := make([]byte, 16)
    if _, err := rand.Read(b); err != nil {
        return "", err
    }
    return hex.EncodeToString(b), nil
}

func agent(w http.ResponseWriter, r *http.Request) {
    traceID := r.Header.Get("Traceparent")
    if traceID == "" {
        var err error
        traceID, err = newTraceID()
        if err != nil {
            http.Error(w, "trace allocation failed", http.StatusInternalServerError)
            return
        }
    }
    w.Header().Set("X-Trace-ID", traceID)
    log.Printf(`{"level":"info","event":"agent_request","trace_id":%q}`, traceID)
    w.Header().Set("Content-Type", "application/json")
    fmt.Fprintf(w, `{"status":"accepted","trace_id":%q}`, traceID)
}

func captureBrowserError(w http.ResponseWriter, r *http.Request) {
    defer r.Body.Close()
    var event browserError
    decoder := json.NewDecoder(http.MaxBytesReader(w, r.Body, 16<<10))
    decoder.DisallowUnknownFields()
    if err := decoder.Decode(&event); err != nil || event.TraceID == "" || event.Message == "" {
        http.Error(w, "invalid error report", http.StatusBadRequest)
        return
    }
    log.Printf(`{"level":"error","event":"browser_error","trace_id":%q,"message":%q,"page":%q}`,
        event.TraceID, event.Message, event.Page)
    w.WriteHeader(http.StatusAccepted)
}

func searchLogs(ctxRequest *http.Request, baseURL, apiKey string) ([]byte, error) {
    client := &http.Client{Timeout: 10 * time.Second}
    url := strings.TrimRight(baseURL, "/") + "/v1/logs/search"
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctxRequest.Context(), http.MethodGet, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctxRequest.Context().Done():
                return nil, ctxRequest.Context().Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("log search failed: status=%d body=%s", resp.StatusCode, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("log search remained rate limited")
}

func internalLogSearch(baseURL, apiKey string) http.HandlerFunc {
    return func(w http.ResponseWriter, r *http.Request) {
        body, err := searchLogs(r, baseURL, apiKey)
        if err != nil {
            http.Error(w, err.Error(), http.StatusBadGateway)
            return
        }
        w.Header().Set("Content-Type", "application/json")
        w.Write(body)
    }
}

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    apiKey := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || apiKey == "" {
        log.Fatal("INFRAI_BASE_URL and INFRAI_API_KEY are required")
    }
    mux := http.NewServeMux()
    mux.HandleFunc("POST /agent/run", agent)
    mux.HandleFunc("POST /client-errors", captureBrowserError)
    mux.HandleFunc("GET /internal/log-search", internalLogSearch(baseURL, apiKey))
    server := &http.Server{
        Addr:              ":8080",
        Handler:           mux,
        ReadHeaderTimeout: 5 * time.Second,
    }
    log.Fatal(server.ListenAndServe())
}

In production, validate an incoming W3C traceparent header instead of accepting arbitrary text, or generate the value at a trusted edge. Return a separate, simple X-Trace-ID if support staff need a copyable identifier. The browser report should contain the message, page, release, timestamp, and correlation value after redaction; stack collection is useful, but without source-map processing a minified stack will remain difficult to interpret.

There is also a retry trap. A browser may resend after a timeout even though the server accepted the first report, so give reports a client-generated event ID and deduplicate them at ingestion. This is less glamorous than a trace visualization, but it prevents one flaky connection from becoming five apparent failures.

Duplicates happen.

The ownership test before choosing a product

The decision is not "which observability vendor is best?" It is which evidence must be available during the dispatch SLO's response window, and how much on-call machinery the platform team will own.

Option Strong fit Boundary that changes the decision Operational posture
Sentry Browser exceptions, source maps, releases, and session replay Cost attribution for an AI loop still needs application metadata and a deliberate join key Managed or self-hosted options; browser instrumentation is a first-class concern
Datadog Logs, APM traces, browser monitoring, and service-wide investigation in one suite Broad adoption can increase instrumentation and governance work across teams Managed platform with integrated query and alerting
Honeycomb High-cardinality event analysis and trace-oriented debugging Frontend crash ergonomics are not the same product emphasis as a dedicated browser error tracker Managed observability centered on events and traces
OpenTelemetry plus your storage backends Vendor-neutral generation and transport of traces, metrics, and logs You still choose, operate, and pay for storage, querying, alerting, and browser diagnostics Maximum control, highest platform ownership
Infrai A plain REST error and log path under one key, with no client SDK version to maintain; AI responses can expose per-call cost, vendor, latency, and request metadata Correlation is manual: there is no distributed tracing query or span tree, browser source-map decoding, replay, or built-in alert delivery Light integration, but polling and any alert dispatcher remain yours

This table is deliberately unfair to anyone seeking a universal winner. There isn't one. Sentry is the sharper default when a minified browser stack must resolve to source and replay is operationally important. Datadog fits an organization already standardizing infrastructure, APM, logs, and real-user monitoring under one control plane. Honeycomb is compelling when engineers need to ask unplanned questions of richly structured events. OpenTelemetry is the sensible instrumentation layer when portability is a roadmap requirement, although the collector is not an incident-management product by itself.

The lighter REST option fits a narrower operating model: backend and API error capture, searchable logs, and manual correlation using stored identifiers. It can receive frontend summaries through your server. Because it has no threshold, phone, SMS, or webhook notification route, a team must poll the query surface and operate its own alert dispatcher; because it has no synthetic check or heartbeat monitor, use a service such as Healthchecks.io for silent "the dispatch job never ran" failures. Those are material on-call costs, not footnotes.

Eight events before the browser retries

Start with the error budget. If the dispatch agent has a service-level objective of 99.9% successful runs, one million monthly runs leave an error budget of 1,000 unsuccessful runs. That arithmetic is illustrative capacity planning, not a claim about any product's uptime. Define whether a run that returns a usable plan after one model retry is successful, degraded, or failed before dashboards begin making the decision for you.

Then size ingestion from events per run rather than from average requests per second. Suppose the application deliberately records one run summary, up to six model-call events, and one terminal error event. At 1,000,000 runs, the upper planning bound is 8,000,000 events per month before browser duplicates and retries. Peak dispatch periods matter more than the monthly average, so load-test the ingestion path at the burst rate and verify what happens under backpressure.

Cost attribution should follow the same hierarchy used for reliability: tenant or business unit, workflow, agent run, call. Record the provider-reported call cost where available, then roll it up; do not infer spend from latency, token guesses, or the presence of an exception. A failed loop may contain successful paid calls, and a successful loop may be wasteful because it retried five times.

Measure the loop.

This is where a superficially simple setup can become expensive to operate. Polling for errors needs a schedule, pagination discipline, a durable cursor, deduplication, and its own dead-man check. If the team cannot commit to testing that path and paging on its failure, buy alert delivery from a platform that owns it.

Two different kinds of silence

Use the backend-first pattern when support mainly needs to correlate a user-visible failure with API logs and AI-call cost, and when manual trace lookup is acceptable. It also works as a transitional architecture: preserve W3C trace context now, then add a trace backend later without changing the join semantics.

Do not use it alone for a consumer-facing browser application where minified stack decoding, release health, breadcrumbs, or session replay materially reduce mean time to recovery. Do not pretend identifier search is distributed tracing. A list of logs with matching text cannot show parent-child timing, missing spans, or the critical path.

Finally, separate active failure from silent absence. Error capture observes something that happened. A missed logistics reconciliation job emits nothing, which is why a heartbeat monitor belongs outside the job and outside the same failure domain. Three tools can be simpler than one improvised platform when each has a crisp responsibility.

The decision rule I would put in the roadmap is blunt: choose a dedicated browser monitor if frontend diagnosis drives the incident; choose a full observability suite or OpenTelemetry-backed stack if cross-service trace exploration drives it; choose the light REST pipeline if backend error capture, AI-call cost attribution, and manual ID correlation satisfy the SLO. Revisit the choice when manual joins consume more response time than the integration originally saved.

Sources

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.