Comparing LLM Models: Oxlo.ai's Perspective
I built a lightweight model router that sends the same prompt to four different Oxlo.ai models, scores the replies, and returns the best one. It helps anyone who wants to stop guessing which model fits a task and start l
I built a lightweight model router that sends the same prompt to four different Oxlo.ai models, scores the replies, and returns the best one. It helps anyone who wants to stop guessing which model fits a task and start letting the models compete on actual output.
What you'll need
- Python 3.10 or newer
- An Oxlo.ai API key from https://portal.oxlo.ai
- The OpenAI SDK:
pip install openai
Step 1: Configure the Oxlo.ai client
I start with a single client pointing to Oxlo.ai. Because the API is fully OpenAI-compatible, the setup is identical to what you already know, only the base_url changes.
from openai import OpenAI
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key="YOUR_OXLO_API_KEY"
)
# Quick sanity check
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": "Say hi"}],
max_tokens=10,
)
print(response.choices[0].message.content)
Step 2: Define the candidate pool
Oxlo.ai hosts reasoning, coding, multilingual, and general-purpose models. I picked four that cover distinct strengths and added short trait labels so the final report stays readable.
CANDIDATES = [
{"id": "llama-3.3-70b", "name": "Llama 3.3 70B", "trait": "general-purpose"},
{"id": "qwen-3-32b", "name": "Qwen 3 32B", "trait": "multilingual reasoning"},
{"id": "deepseek-v3.2", "name": "DeepSeek V3.2", "trait": "coding and reasoning"},
{"id": "kimi-k2.6", "name": "Kimi K2.6", "trait": "agentic coding and vision"},
]
Step 3: Query every candidate in parallel
Running them concurrently keeps latency low. I use ThreadPoolExecutor to fire off four requests at once. Oxlo.ai has no cold starts on popular models, so the responses come back immediately.
from concurrent.futures import ThreadPoolExecutor, as_completed
def ask_model(candidate, user_prompt):
response = client.chat.completions.create(
model=candidate["id"],
messages=[{"role": "user", "content": user_prompt}],
temperature=0.2,
max_tokens=512,
)
return {
"candidate": candidate,
"text": response.choices[0].message.content.strip(),
}
def gather_responses(user_prompt):
results = []
with ThreadPoolExecutor(max_workers=4) as exe:
futures = {exe.submit(ask_model, c, user_prompt): c for c in CANDIDATES}
for future in as_completed(futures):
results.append(future.result())
return results
Step 4: Build the judge prompt
I use a separate judge call to score the outputs. The judge is Llama 3.3 70B because it follows rubrics reliably. Keeping the system prompt strict makes the scoring consistent.
JUDGE_SYSTEM_PROMPT = """You are an expert evaluator. You will receive a user prompt and several candidate answers from different LLMs.
Score each answer on three criteria from 1 to 5:
1. Accuracy: is the answer correct and well-reasoned?
2. Clarity: is the answer easy to read and unambiguous?
3. Completeness: does it cover all parts of the prompt?
Respond in this exact format for each candidate:
Model: [name]
Accuracy: [int]
Clarity: [int]
Completeness: [int]
Total: [sum]
Then declare a winner.
Do not add extra commentary outside the requested format."""
def judge_responses(user_prompt, results):
block = f"User prompt: {user_prompt}\n\n"
for r in results:
block += f"--- {r['candidate']['name']} ---\n{r['text']}\n\n"
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": JUDGE_SYSTEM_PROMPT},
{"role": "user", "content": block},
],
temperature=0.1,
max_tokens=1024,
)
return response.choices[0].message.content
Step 5: Assemble the router
This ties the pieces together. It gathers responses, runs the judge, and returns the raw outputs plus the evaluation. Because Oxlo.ai charges a flat rate per request, the cost of this four-way comparison is just four predictable requests, regardless of how long the prompt is.
def compare_and_route(user_prompt):
print(f"Running comparison for: {user_prompt[:60]}...\n")
results = gather_responses(user_prompt)
verdict = judge_responses(user_prompt, results)
print("=== Raw outputs ===")
for r in results:
label = f"{r['candidate']['name']} ({r['candidate']['trait']})"
print(f"\n{label}\n{r['text'][:300]}...")
print("\n=== Judge verdict ===")
print(verdict)
return results, verdict
Run it
I tested this on a prompt that mixes reasoning and code. Call the router from a small entry block like this.
if __name__ == "__main__":
PROMPT = (
"Write a Python function that checks if a string is a palindrome, "
"ignoring spaces and punctuation. Explain the time complexity."
)
compare_and_route(PROMPT)
Example output:
Running comparison for: Write a Python function that checks if a string is a palindrome...
=== Raw outputs ===
Llama 3.3 70B (general-purpose)
Here is a Python function... O(n) time...
Qwen 3 32B (multilingual reasoning)
```python
def is_palindrome(s):...
DeepSeek V3.2 (coding and reasoning)
You can solve this with two pointers... O(n)...
Kimi K2.6 (agentic coding and vision)
```
python
def clean_and_check(text):...
=== Judge verdict ===
Model: Llama 3.3 70B
Accuracy: 5
Clarity: 5
Completeness: 4
Total: 14
Model: DeepSeek V3.2
Accuracy: 5
Clarity: 5
Completeness: 5
Total: 15
Winner: DeepSeek V3.2
Wrap-up and next steps
You can extend this into a permanent router by caching the winner per prompt type, or swap the judge model for Kimi K2.6 when you need vision-aware evaluation. If you want to see how the flat per-request pricing keeps costs predictable even when you send long prompts to four models at once, check the details at https://oxlo.ai/pricing.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.