Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Evaluating LLMs: Metrics and Best Practices

We are going to build a self-contained evaluation harness that scores answers from any LLM against a ground-truth dataset. It combines an LLM-as-a-judge scorer running on Oxlo.ai with classic string-similarity metrics, g

We are going to build a self-contained evaluation harness that scores answers from any LLM against a ground-truth dataset. It combines an LLM-as-a-judge scorer running on Oxlo.ai with classic string-similarity metrics, giving you a repeatable baseline for model selection.

What you'll need

Oxlo.ai works as a drop-in replacement for the OpenAI client, so the same SDK handles authentication and requests.

Step 1: Prepare the benchmark dataset

Hardcode a small JSONL-style dataset of three questions with reference answers so the script is fully reproducible without external files.

BENCHMARK = [
    {
        "question": "What is the capital of France?",
        "reference": "The capital of France is Paris."
    },
    {
        "question": "Explain the difference between a list and a tuple in Python.",
        "reference": "A list is mutable, meaning it can be changed after creation, while a tuple is immutable and cannot be modified."
    },
    {
        "question": "Who wrote '1984'?",
        "reference": "George Orwell wrote the novel '1984', published in 1949."
    }
]

Step 2: Generate candidate answers

Query Oxlo.ai using Llama 3.3 70B to generate answers for every question. We use the OpenAI SDK pointed at the Oxlo.ai base URL.

from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key=os.environ.get("OXLO_API_KEY")
)

def generate_answer(question: str) -> str:
    response = client.chat.completions.create(
        model="llama-3.3-70b",
        messages=[
            {"role": "system", "content": "Answer the question concisely and accurately."},
            {"role": "user", "content": question},
        ],
        temperature=0.2,
        max_tokens=256,
    )
    return response.choices[0].message.content.strip()

candidates = []
for item in BENCHMARK:
    answer = generate_answer(item["question"])
    candidates.append({
        "question": item["question"],
        "reference": item["reference"],
        "candidate": answer
    })
    print(f"Q: {item['question']}\nA: {answer}\n")

Step 3: Define the judge prompt

We use Kimi K2.6 on Oxlo.ai as the judge. The system prompt asks it to rate correctness and relevance on a 1 to 5 scale and return strict JSON.

JUDGE_SYSTEM_PROMPT = """You are an expert evaluator. Compare the candidate answer to the reference answer.

Score the candidate on two metrics from 1 to 5:
- correctness: factual alignment with the reference.
- relevance: whether the answer addresses the question.

Respond with ONLY a JSON object in this exact format:
{"correctness": , "relevance": , "reason": ""}
"""

Step 4: Run the LLM judge

For each candidate answer, call the judge model and parse the JSON response. We use json.loads with a regex fallback to strip markdown fences if the model emits them.

import json
import re

def judge_answer(question: str, reference: str, candidate: str) -> dict:
    user_prompt = f"""Question: {question}
Reference Answer: {reference}
Candidate Answer: {candidate}

Provide your JSON evaluation now."""

    response = client.chat.completions.create(
        model="kimi-k2.6",
        messages=[
            {"role": "system", "content": JUDGE_SYSTEM_PROMPT},
            {"role": "user", "content": user_prompt},
        ],
        temperature=0.1,
        max_tokens=256,
    )
    raw = response.choices[0].message.content.strip()

    # Extract JSON if wrapped in markdown fences
    match = re.search(r'\{.*\}', raw, re.DOTALL)
    if not match:
        raise ValueError(f"Could not parse judge output: {raw}")
    return json.loads(match.group(0))

for item in candidates:
    scores = judge_answer(item["question"], item["reference"], item["candidate"])
    item["llm_scores"] = scores
    print(scores)

Step 5: Compute code-based metrics

Add deterministic metrics using Python's difflib. Exact match is binary, and sequence ratio gives a continuous similarity score.

from difflib import SequenceMatcher

def compute_metrics(reference: str, candidate: str) -> dict:
    exact = 1.0 if reference.strip().lower() == candidate.strip().lower() else 0.0
    similarity = SequenceMatcher(None, reference, candidate).ratio()
    return {
        "exact_match": round(exact, 2),
        "sequence_similarity": round(similarity, 3),
    }

for item in candidates:
    metrics = compute_metrics(item["reference"], item["candidate"])
    item["code_metrics"] = metrics
    print(metrics)

Step 6: Aggregate and report

Combine LLM and code metrics into a final report, then print a summary table and save the raw results to a JSON file.

def build_report(results: list) -> None:
    print(f"{'Question':<45} {'Correctness':>12} {'Relevance':>10} {'Exact':>6} {'Sim':>6}")
    print("-" * 85)
    for r in results:
        q = r["question"][:44]
        c = r["llm_scores"]["correctness"]
        rel = r["llm_scores"]["relevance"]
        ex = r["code_metrics"]["exact_match"]
        sim = r["code_metrics"]["sequence_similarity"]
        print(f"{q:<45} {c:>12} {rel:>10} {ex:>6} {sim:>6}")

    with open("eval_results.json", "w") as f:
        json.dump(results, f, indent=2)
    print("\nFull results written to eval_results.json")

build_report(candidates)

Run it

Export your key and execute the script. The output below shows what a typical run looks like when the candidate model answers correctly but not verbatim.

$ export OXLO_API_KEY="sk-oxlo.ai-..."
$ python eval_harness.py

Q: What is the capital of France?
A: Paris is the capital of France.

Q: Explain the difference between a list and a tuple in Python.
A: Lists are mutable, whereas tuples are immutable.

Q: Who wrote '1984'?
A: George Orwell wrote the novel '1984'.

{'correctness': 5, 'relevance': 5, 'reason': 'Factually correct and directly answers the question.'}
{'correctness': 5, 'relevance': 5, 'reason': 'Captures the core distinction accurately.'}
{'correctness': 5, 'relevance': 5, 'reason': 'Accurately identifies the author.'}

{'exact_match': 0.0, 'sequence_similarity': 0.612}
{'exact_match': 0.0, 'sequence_similarity': 0.485}
{'exact_match': 0.0, 'sequence_similarity': 0.734}

Question                                      Correctness  Relevance  Exact    Sim
-------------------------------------------------------------------------------------
What is the capital of France?                            5          5    0.0  0.612
Explain the difference between a list and a t             5          5    0.0  0.485
Who wrote '1984'?                                         5          5    0.0  0.734

Full results written to eval_results.json

Wrap-up

You now have a reproducible harness that mixes deterministic metrics with an LLM judge. Two concrete next steps: plug in a different candidate model such as Qwen 3 32B or DeepSeek V3.2 on Oxlo.ai to A/B test quality, or extend the judge prompt to evaluate faithfulness against a retrieved context block for RAG pipelines.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.