Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Route Once, Fail Over Among Equals: When NOT to Retry an LLM Call

Most LLM APIs I review have one IChatClient and a prayer. Hello goes to the same model as "why does this async code deadlock under load." When that provider returns 503 for twenty minutes, so does the product. Routing a

Most LLM APIs I review have one IChatClient and a prayer. Hello goes to the same model as "why does this async code deadlock under load." When that provider returns 503 for twenty minutes, so does the product.

Routing and failover both matter. Mixing them up is how you get a green uptime chart and worse answers. The full sample (Microsoft.Extensions.AI routing types, circuit breaker, tests, offline profile) is in Tech Skill Builder. Here is the short version: a few hard rules, then a Q&A on when not to fail over.

The one architecture rule that saves you

Pick the route once. Fail over only among equals.

SemanticRoutingChatClient          // meaning → route "deep" or "fast"
  └─ CircuitBreakingFailoverChatClient
       ├─ model A (timeout wrapper)
       └─ model B (backup, maybe another vendor)

SemanticRoutingChatClient embeds the prompt and scores it against example utterances per route. Failover then walks that route's ordered model list. Invert the layers and an outage on your reasoning model can quietly dump a hard debugging question onto a small-talk model. Uptime stays green. Quality does not.

The MEAI routing and failover types are marked [Experimental("MEAI001")] in 10.x. Opt in once (for example in Directory.Build.props) so the choice is visible.

Config shape: Connections / Models / Routes

Keep secrets out of policy:

Section Holds
Connections endpoints and keys (OpenAI, Azure, Ollama, Offline)
Models connection + model/deployment name + per-attempt timeout
Routes ordered models, utterances, optional instructions

Rotating a key touches one connection. Swapping a model touches one entry. Routing policy never contains a secret, so you can review it in a PR without redacting.

Validate on start. Unknown model names, a missing default route, or a live connection still set to YOUR_OPENAI_API_KEY should fail the deployment, not the first request after lunch.

builder.Services
    .AddOptions<AiRoutingOptions>()
    .BindConfiguration(AiRoutingOptions.SectionName)
    .ValidateDataAnnotations() // or IValidateOptions
    .ValidateOnStart();
// full graph checks (route → models → connections) in the complete project

Q&A: when should you NOT fail over?

Q: The provider returned 400. Try the backup?

No. A malformed request or schema mistake will fail the same way on the next model. You pay twice for the same bug.

Q: Content filter / policy rejection?

No. The next vendor often rejects the same prompt. Treat it as a final answer for this request (and a product/safety signal), not an outage.

Q: 401 / bad API key?

No. Failing over to another model on the same broken credential burns time. Fix the secret. Cross-vendor backups only help if their credentials are healthy.

Q: 429, 5xx, timeout, connection reset?

Yes. Those are the transient cases. Fail over. Prefer a per-attempt timeout that surfaces as TimeoutException (not a naked OperationCanceledException from the request abort) so policy can classify it.

Skeleton of the classifier idea:

static bool IsTransient(Exception ex) => ex switch
{
    TimeoutException => true,
    HttpRequestException { StatusCode: null } => true, // DNS, reset, …
    HttpRequestException { StatusCode: { } s }
        when (int)s is 408 or 429 or >= 500 => true,
    // ClientResultException status 0 / 408 / 429 / 5xx → true
    _ => false, // 400, 401, content filter, etc. → stop the chain
};
// full implementation in the complete project

Q: We already streamed the first token. Can we switch providers?

No. FailoverChatClient commits once the first update reaches the caller. You cannot un-send half an answer. After commitment, a failure is final for that stream. Design UX accordingly (partial text + clear failed path, or non-streaming for critical paths).

Q: The same model has timed out three times in a row. Keep trying it first?

Not on every request. Open a circuit for a break duration, skip that model, probe later. Otherwise every call pays a full timeout tax while the provider is known-bad.

// CircuitBreakingFailoverChatClient : FailoverChatClient
protected override ValueTask<IChatClient> SelectClientAsync(...)
{
    // prefer first untried model with a closed/half-open circuit
    // if all open, still probe one untried candidate (slow > guaranteed 503)
    // full implementation in the complete project
}

Q: Embeddings for routing are down. Should the whole API die?

Prefer a default route. Semantic routing is an optimization. When embedding fails, fall back to DefaultRoute (often fast) so chat still answers. Log loudly; do not pretend routing ran.

Checklist before you ship

  • [ ] Router outside failover (route once per request)
  • [ ] Failover list = interchangeable models for that route
  • [ ] Transient vs non-transient classifier (do not retry 400 / filter / auth)
  • [ ] Respect streaming commitment
  • [ ] Per-model timeout + circuit breaker
  • [ ] Connections / Models / Routes split + ValidateOnStart
  • [ ] Response includes route, model, and attempt trace (so you can debug without guessing)
  • [ ] Offline or local backup on the chain if you want "always answer something"

What this post skips

Wiring OrderedFailoverChatClient vs a custom FailoverChatClient subclass, fault-injection endpoints for demos, and the full options validator live in the complete project. Use this as the decision guide; use the repo when you want something you can dotnet test tonight.

If you are wiring multi-model chat in .NET, the ready-to-run package is in Tech Skill Builder: semantic routing, circuit-breaking failover, config validation, and an offline profile so you can practice outages without burning API credits. Limited-time membership pricing is on the product page. Grab it once, then spend your time on prompts and product, not on reinventing which errors deserve a second try.

Uptime without quality is a lie. Route for meaning, fail over among equals, and refuse to retry mistakes that will not heal.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.