Dev.to AI 🤖 Ai 👁 0 📖 1 min read

Why I put one thin layer between my app and every LLM provider

On my first AI feature I called one provider straight from the code that needed it. It worked, and for a while that was the right call. Then I wanted a cheaper model for the boring tasks, like tagging and short summarie

On my first AI feature I called one provider straight from the code that needed it. It worked, and for a while that was the right call.

Then I wanted a cheaper model for the boring tasks, like tagging and short summaries. That's when I noticed the provider was everywhere: keys in three places, retry logic copied around, and no clear idea what each feature actually cost.

Chalk's post this week cited an a16z survey where 81% of CIOs at Global 2000 companies now use three or more model families. I'm nowhere near that scale, but even with two providers the mess shows up fast.

What the layer does

It's small. Roughly:

  • One function per task type, like summarize or classify, not per provider.
  • The model for each task comes from config, so switching is a one line change.
  • A timeout and one fallback model, so a slow provider doesn't hang a user request.
  • A daily spend limit per key. I got this one after a retry loop ran all night once.
  • A log line per request with task, model, tokens and latency.

What it cost me

Another thing to maintain. Tool calling and structured output formats differ between providers, so the abstraction leaks. I ended up with small adapter code per provider anyway.

It also tempts you to over build. I'd skip it until a second provider is actually on the table.

What I'd do again

The per task log. Seeing that one feature used most of the tokens changed what I worked on next more than any routing trick did.

If you're running more than one model, did you build your own layer or use a gateway?

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.