One API Key Across Model Providers: Compare Token Cost for Healthtech Moderation
Send every moderation report through one server-side gateway, but don't choose the model by the cheapest token rate. Short answer: one API key across OpenAI, Claude, and Gemini is useful for startup operations; it isn't
Send every moderation report through one server-side gateway, but don't choose the model by the cheapest token rate. Short answer: one API key across OpenAI, Claude, and Gemini is useful for startup operations; it isn't a sound routing rule. For a healthtech app, compare total cost per accepted classification, validate every response, and reserve slower inference for ambiguous cases before a human reviews them.
That distinction matters in a one-person SaaS. A low token invoice is irrelevant if weak classifications create more review work or if a slow first pass leaves the queue idle. I want the undifferentiated authentication and transport work behind one adapter so I can ship weekly. The business logic stays mine.
How should one API key compare token cost across providers?
The job is narrow: classify a moderation report before human review. A useful output might contain one label, a bounded confidence value, and a short reason. It must never make the final moderation decision. The reviewer owns that action.
Start there.
The cheapest advertised input rate cannot describe the cost of this workflow. Reports vary in length, providers count tokens under their own published pricing terms, outputs vary, and malformed responses consume retries or reviewer attention. A router that presents one key across OpenAI, Claude, and Gemini still has to preserve those differences in its usage records. Compare candidates on a fixed evaluation set instead. Record input tokens, output tokens, end-to-end latency, schema validity, and agreement with labels approved by the moderation team. For a startup app, that accepted-result record is more useful than a copied price sheet: it connects spend to work that actually reached the review queue, while leaving room to replace an upstream provider without rewriting the job.
I would use four labels for the first build: urgent_safety, policy_review, duplicate, and insufficient_context. Four is a product constraint, not a universal taxonomy. It keeps the queue useful while the team can still inspect every category. Add labels only when they change what a reviewer does.
The decision table is small on purpose:
| Signal | Gate | Why it matters |
|---|---|---|
| Schema validity | Must pass | Invalid data cannot enter the review queue |
| Critical-label recall | Set with the clinical and safety owners | Missing an urgent report has asymmetric impact |
| p95 end-to-end latency | Fit the review workflow | Fast tokens do not guarantee a timely classification |
| Cost per accepted result | Compare after retries | Raw token cost hides failed attempts |
| Reviewer override rate | Watch by label | Drift appears as human rework |
Do not collapse those columns into one magic score too early. Quality is a gate. Latency and cost are optimization targets after that gate passes.
No magic score.
The smallest working gateway
The gateway needs one public interface, a provider adapter per backend, and no provider-specific objects outside those adapters. Keep its credential on the server. A browser or mobile client should submit a report to your application, never receive the upstream key.
This TypeScript sketch makes the boundary concrete. It assumes each adapter has already translated its provider's structured-output mechanism into the common result. No route names, model IDs, or changing prices are baked into the domain code.
type Label =
| "urgent_safety"
| "policy_review"
| "duplicate"
| "insufficient_context";
type Classification = {
label: Label;
confidence: number;
reason: string;
};
type Candidate = "fast" | "careful";
type Runtime = {
classify(candidate: Candidate, report: string): Promise<Classification>;
};
const labels = new Set<Label>([
"urgent_safety",
"policy_review",
"duplicate",
"insufficient_context",
]);
function validate(value: Classification): Classification {
if (!labels.has(value.label)) throw new Error("invalid label");
if (!Number.isFinite(value.confidence)) throw new Error("invalid confidence");
if (value.confidence < 0 || value.confidence > 1) {
throw new Error("confidence outside 0..1");
}
if (value.reason.trim().length === 0) throw new Error("missing reason");
return value;
}
export async function triage(
runtime: Runtime,
report: string,
): Promise<Classification> {
const first = validate(await runtime.classify("fast", report));
// The threshold is calibrated on reviewed data, not copied from a model card.
if (first.confidence >= 0.82 && first.label !== "insufficient_context") {
return first;
}
return validate(await runtime.classify("careful", report));
}
The 0.82 threshold is deliberately visible. It is an example policy value to calibrate, not a claim about model reliability. Store its version beside each result. Otherwise a threshold change can look like model drift during an audit.
Call the first policy triage-v1. Names beat guesswork six months later.
Structured output support can reduce parsing ambiguity, but application-side validation still belongs at the boundary. OpenAI's Structured Outputs guide documents schema-constrained responses for that API; other adapters must meet the same internal contract through their supported mechanisms. The contract is the asset, not any one provider's request shape.
Measure accepted work, not cheap tokens
Run the same frozen, de-identified set through every candidate. Split it by label and by input-length band so a good average cannot hide a weak safety category. Keep the human-approved label, but also preserve disagreement notes; moderation language can be ambiguous, and a bare gold label can conceal that ambiguity.
Then calculate cost from recorded usage under the billing terms that applied during the run. Avoid hard-coding a public price table into the router. Prices and model catalogs change, while your stored usage and invoice remain the evidence for that period.
A useful denominator is accepted classifications: outputs that pass the schema and the quality gate without another model call. Track retries separately. Also measure wall-clock time from gateway receipt to validated result, because network time, queueing, inference, and fallback all reach the reviewer together.
Retries aren't free.
This changes the apparent winner. A candidate with lower token cost but frequent escalation can cost more per accepted result and increase p95 latency. A more capable candidate may be wasteful on obvious duplicates. Route by observed workload segment, then rerun the evaluation whenever the prompt, taxonomy, adapter, or candidate model changes.
Keep health data out of prompts unless the workflow and provider arrangement explicitly permit it. For evaluation, use de-identified examples approved for that purpose, minimize retained text, and log identifiers and metrics rather than raw reports wherever possible. The exact compliance controls depend on the system's jurisdiction and agreements; a generic router cannot decide them.
Failure handling belongs in the product
Timeouts, invalid JSON, unknown labels, and exhausted retries should produce an explicit needs_human_review state in the application. They should not silently become policy_review, because that mixes infrastructure failure with a model judgment. Keep the original report available to the authorized reviewer and attach the attempt metadata.
Streaming rarely improves this classification job. The application needs one small validated object, so partial tokens add state without letting the reviewer act sooner. Server-Sent Events are a one-way server-to-client channel and use the text/event-stream format; they make sense if the product needs live queue progress, not as a requirement for model routing. MDN documents the browser behavior and connection considerations.
Use bounded timeouts and a strict retry budget. One retry may handle a transient transport failure; repeated semantic failures should go to review rather than multiply spend and delay. Log candidate ID, policy version, latency, usage reported by the adapter, validation outcome, and fallback reason. Never log the shared credential.
The gateway earns its keep when a provider change is an adapter edit and an evaluation run, not a rewrite of moderation policy. It does not remove provider differences. It makes them measurable.
What I would change at scale
At low volume, an in-process policy and a database table are enough. Outsource the generic transport layer if its contract fits, but retain evaluation fixtures, routing thresholds, and audit records. Those are specific to the product and determine whether automation helps the reviewer.
At higher volume, move classification to a queue, make requests idempotent, and separate the latency objective for urgent reports from the bulk objective for duplicates. Add a shadow lane before changing the active candidate: send an approved sample to the challenger, compare it offline, and do not let the shadow response affect the live queue.
The trade-off is operational weight. Queues, shadow traffic, and per-segment policies create more states to observe. Add them only after measurements show that the simple two-stage path misses a quality or latency target. Shipping weekly means refusing infrastructure that has no measured job yet.
One key is convenient. The durable design is one internal contract, versioned policy, and evidence for every routing change. That lets a small team compare providers without turning mutable token prices into product logic or outsourcing the safety decision.
References
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.