Evaluating LLMs: Metrics and Best Practices
We are going to build a self-contained evaluation harness that scores answers from any LLM against a ground-truth dataset. It combines an LLM-as-a-judge scorer running on Oxlo.ai with classic string-similarity metrics, g
We are going to build a self-contained evaluation harness that scores answers from any LLM against a ground-truth dataset. It combines an LLM-as-a-judge scorer running on Oxlo.ai with classic string-similarity metrics, giving you a repeatable baseline for model selection.
What you'll need
- Python 3.10 or newer
- The OpenAI SDK:
pip install openai - An Oxlo.ai API key from https://portal.oxlo.ai
Oxlo.ai works as a drop-in replacement for the OpenAI client, so the same SDK handles authentication and requests.
Step 1: Prepare the benchmark dataset
Hardcode a small JSONL-style dataset of three questions with reference answers so the script is fully reproducible without external files.
BENCHMARK = [
{
"question": "What is the capital of France?",
"reference": "The capital of France is Paris."
},
{
"question": "Explain the difference between a list and a tuple in Python.",
"reference": "A list is mutable, meaning it can be changed after creation, while a tuple is immutable and cannot be modified."
},
{
"question": "Who wrote '1984'?",
"reference": "George Orwell wrote the novel '1984', published in 1949."
}
]
Step 2: Generate candidate answers
Query Oxlo.ai using Llama 3.3 70B to generate answers for every question. We use the OpenAI SDK pointed at the Oxlo.ai base URL.
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.environ.get("OXLO_API_KEY")
)
def generate_answer(question: str) -> str:
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": "Answer the question concisely and accurately."},
{"role": "user", "content": question},
],
temperature=0.2,
max_tokens=256,
)
return response.choices[0].message.content.strip()
candidates = []
for item in BENCHMARK:
answer = generate_answer(item["question"])
candidates.append({
"question": item["question"],
"reference": item["reference"],
"candidate": answer
})
print(f"Q: {item['question']}\nA: {answer}\n")
Step 3: Define the judge prompt
We use Kimi K2.6 on Oxlo.ai as the judge. The system prompt asks it to rate correctness and relevance on a 1 to 5 scale and return strict JSON.
JUDGE_SYSTEM_PROMPT = """You are an expert evaluator. Compare the candidate answer to the reference answer.
Score the candidate on two metrics from 1 to 5:
- correctness: factual alignment with the reference.
- relevance: whether the answer addresses the question.
Respond with ONLY a JSON object in this exact format:
{"correctness": , "relevance": , "reason": ""}
"""
Step 4: Run the LLM judge
For each candidate answer, call the judge model and parse the JSON response. We use json.loads with a regex fallback to strip markdown fences if the model emits them.
import json
import re
def judge_answer(question: str, reference: str, candidate: str) -> dict:
user_prompt = f"""Question: {question}
Reference Answer: {reference}
Candidate Answer: {candidate}
Provide your JSON evaluation now."""
response = client.chat.completions.create(
model="kimi-k2.6",
messages=[
{"role": "system", "content": JUDGE_SYSTEM_PROMPT},
{"role": "user", "content": user_prompt},
],
temperature=0.1,
max_tokens=256,
)
raw = response.choices[0].message.content.strip()
# Extract JSON if wrapped in markdown fences
match = re.search(r'\{.*\}', raw, re.DOTALL)
if not match:
raise ValueError(f"Could not parse judge output: {raw}")
return json.loads(match.group(0))
for item in candidates:
scores = judge_answer(item["question"], item["reference"], item["candidate"])
item["llm_scores"] = scores
print(scores)
Step 5: Compute code-based metrics
Add deterministic metrics using Python's difflib. Exact match is binary, and sequence ratio gives a continuous similarity score.
from difflib import SequenceMatcher
def compute_metrics(reference: str, candidate: str) -> dict:
exact = 1.0 if reference.strip().lower() == candidate.strip().lower() else 0.0
similarity = SequenceMatcher(None, reference, candidate).ratio()
return {
"exact_match": round(exact, 2),
"sequence_similarity": round(similarity, 3),
}
for item in candidates:
metrics = compute_metrics(item["reference"], item["candidate"])
item["code_metrics"] = metrics
print(metrics)
Step 6: Aggregate and report
Combine LLM and code metrics into a final report, then print a summary table and save the raw results to a JSON file.
def build_report(results: list) -> None:
print(f"{'Question':<45} {'Correctness':>12} {'Relevance':>10} {'Exact':>6} {'Sim':>6}")
print("-" * 85)
for r in results:
q = r["question"][:44]
c = r["llm_scores"]["correctness"]
rel = r["llm_scores"]["relevance"]
ex = r["code_metrics"]["exact_match"]
sim = r["code_metrics"]["sequence_similarity"]
print(f"{q:<45} {c:>12} {rel:>10} {ex:>6} {sim:>6}")
with open("eval_results.json", "w") as f:
json.dump(results, f, indent=2)
print("\nFull results written to eval_results.json")
build_report(candidates)
Run it
Export your key and execute the script. The output below shows what a typical run looks like when the candidate model answers correctly but not verbatim.
$ export OXLO_API_KEY="sk-oxlo.ai-..."
$ python eval_harness.py
Q: What is the capital of France?
A: Paris is the capital of France.
Q: Explain the difference between a list and a tuple in Python.
A: Lists are mutable, whereas tuples are immutable.
Q: Who wrote '1984'?
A: George Orwell wrote the novel '1984'.
{'correctness': 5, 'relevance': 5, 'reason': 'Factually correct and directly answers the question.'}
{'correctness': 5, 'relevance': 5, 'reason': 'Captures the core distinction accurately.'}
{'correctness': 5, 'relevance': 5, 'reason': 'Accurately identifies the author.'}
{'exact_match': 0.0, 'sequence_similarity': 0.612}
{'exact_match': 0.0, 'sequence_similarity': 0.485}
{'exact_match': 0.0, 'sequence_similarity': 0.734}
Question Correctness Relevance Exact Sim
-------------------------------------------------------------------------------------
What is the capital of France? 5 5 0.0 0.612
Explain the difference between a list and a t 5 5 0.0 0.485
Who wrote '1984'? 5 5 0.0 0.734
Full results written to eval_results.json
Wrap-up
You now have a reproducible harness that mixes deterministic metrics with an LLM judge. Two concrete next steps: plug in a different candidate model such as Qwen 3 32B or DeepSeek V3.2 on Oxlo.ai to A/B test quality, or extend the judge prompt to evaluate faithfulness against a retrieved context block for RAG pipelines.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.