Best Practices for LLM Development
Building production-grade applications with large language models requires more than wrapping a chat endpoint. You need composable prompts, structured outputs, managed context, and failure modes that degrade gracefully.
Building production-grade applications with large language models requires more than wrapping a chat endpoint. You need composable prompts, structured outputs, managed context, and failure modes that degrade gracefully. The infrastructure choices you make early, including how you are billed for inference, directly affect your architecture. This guide covers practical patterns for reliable LLM development, with concrete examples you can apply today.
Version and Modularize Your Prompts
Treat prompts as source code. Store them in version control instead of embedding them as raw strings in application logic. Use a templating engine such as Jinja2 to inject variables and compose reusable fragments. Break complex tasks into smaller, testable steps rather than monolithic zero-shot prompts. This separation makes A/B testing, regression detection, and prompt collaboration straightforward.
Enforce Structured Outputs with JSON Mode
Unstructured text is fragile. Use JSON mode to guarantee parseable responses, and validate outputs against a schema before they reach business logic. Oxlo.ai supports JSON mode across its chat models through a fully OpenAI-compatible API, so you can rely on consistent formatting without adopting a vendor-specific SDK.
from openai import OpenAI
from pydantic import BaseModel, Field
import os
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.getenv("OXLO_API_KEY")
)
class SentimentResult(BaseModel):
sentiment: str = Field(pattern="^(positive|negative|neutral)$")
confidence: float = Field(ge=0.0, le=1.0)
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": "Analyze sentiment. Return only JSON."},
{"role": "user", "content": "The API latency is excellent and documentation is clear."}
],
response_format={"type": "json_object"}
)
result = SentimentResult.model_validate_json(response.choices[0].message.content)
Manage Context Windows Efficiently
Long prompts degrade latency and, on token-based platforms, inflate costs. Use retrieval to inject only relevant context, and summarize earlier conversation turns when possible. For agentic workflows that accumulate tool results and history, context length grows quickly. On Oxlo.ai, request-based pricing means your cost stays flat per call regardless of how much context you pass, so you can prioritize accuracy over token economy. This shifts the architectural tradeoff from cost compression to relevance.
Design Tool Use for Deterministic Execution
Function calling is the bridge between probabilistic models and deterministic systems. Keep function descriptions explicit and parameters minimal. Always validate arguments server-side before execution, because the model output is not a guarantee of safety. Oxlo.ai supports function calling and tool use across its chat and reasoning models, including Qwen 3 32B and Kimi K2.6.
tools = [
{
"type": "function",
"function": {
"name": "query_database",
"description": "Execute a read-only SQL query",
"parameters": {
"type": "object",
"properties": {
"sql": {"type": "string", "description": "A valid SELECT statement"}
},
"required": ["sql"]
}
}
}
]
response = client.chat.completions.create(
model="qwen3-32b",
messages=[{"role": "user", "content": "How many users signed up last week?"}],
tools=tools,
tool_choice="auto"
)
if response.choices[0].message.tool_calls:
call = response.choices[0].message.tool_calls[0]
# Validate and execute call.function.arguments here
Select Models by Task, Not Benchmark Hype
Do not default to the largest model for every call. Use smaller specialized models for routing, classification, and extraction. Reserve large reasoning models for complex coding or multi-step planning. Oxlo.ai hosts over 45 models across seven categories, from code specialists like Qwen 3 Coder 30B and DeepSeek V3.2 to vision models like Kimi VL A3B. Routing requests to the right model tier is easier when your pricing is predictable per request.
Handle Failures with Streaming and Retries
Networks stall and models occasionally refuse or hallucinate. Implement exponential backoff with jitter. Stream responses to improve perceived latency and allow early cancellation. Oxlo.ai serves popular models with no cold starts, which removes a common source of timeout errors during traffic spikes.
import time
for attempt in range(3):
try:
stream = client.chat.completions.create(
model="deepseek-r1-671b",
messages=[{"role": "user", "content": "Refactor this function to use async/await."}],
stream=True,
timeout=30
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
break
except Exception:
time.sleep(2 ** attempt)
Evaluate Continuously with Real and Synthetic Data
Unit test your prompts with fixed inputs and expected outputs. Use an LLM as a judge only for subjective dimensions like tone or helpfulness, and ground factual correctness against structured assertions. Run regression suites before deploying prompt or model changes. You can use a strong reasoning model such as DeepSeek R1 671B or GLM 5 on Oxlo.ai as a judge without worrying about per-token judge costs, because flat per-request pricing keeps evaluation batches predictable.
Control Costs with Request-Based Pricing
Token-based billing creates tension between context quality and cost. Developers strip documentation, truncate history, or avoid agentic loops to save money. Oxlo.ai uses request-based pricing: one flat cost per API call, regardless of prompt length. For long-context retrieval, multi-turn agents, and batch evaluation, this is often significantly cheaper than scaling costs with every token. You can view the current plans at https://oxlo.ai/pricing.
Conclusion
Reliable LLM development is an engineering discipline. Version your prompts, enforce schemas, handle failures, and choose infrastructure that aligns with how you actually build. Oxlo.ai provides an OpenAI-compatible API with over 45 models, flat per-request pricing, and no cold starts, making it a natural fit for teams shipping agentic and long-context applications.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.