LLM for Text Classification: Best Practices and Use Cases
We are going to build a production-ready support ticket classifier that reads raw customer messages, assigns a category and priority, and routes low-confidence cases to human review. This cuts first-response time and kee
We are going to build a production-ready support ticket classifier that reads raw customer messages, assigns a category and priority, and routes low-confidence cases to human review. This cuts first-response time and keeps support agents focused on issues that actually need a human. I am using Oxlo.ai because its request-based pricing means a ten-word ticket costs the same as a ten-thousand-word transcript, which keeps batch-processing costs flat no matter how verbose your inbound text is.
What you'll need
- Python 3.10 or newer
- The OpenAI SDK:
pip install openai - An Oxlo.ai API key from https://portal.oxlo.ai. Oxlo.ai is fully OpenAI SDK compatible, so the client setup is a single base-URL change.
Step 1: Configure the client and verify the endpoint
I start by instantiating the OpenAI client with the Oxlo.ai base URL. A quick ping confirms the key is live and that there are no cold starts to warm through.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.environ["OXLO_API_KEY"]
)
# Verify connectivity with a cheap call.
ping = client.chat.completions.create(
model="deepseek-v3.2",
messages=[{"role": "user", "content": "ping"}],
max_tokens=5
)
print(ping.choices[0].message.content)
Step 2: Design the classification schema and system prompt
The system prompt is the contract. I keep the category list short, force JSON output, and add a confidence score so I can filter later.
SYSTEM_PROMPT = """You are a support-triage classifier.
Read the user's message and return a single JSON object with these exact keys:
category: one of [Billing, Technical, Account, Feature_Request, General]
priority: one of [low, medium, high, critical]
sentiment: one of [positive, neutral, negative]
summary: 5 to 10 words describing the core issue
confidence: integer 1-10 representing classification certainty
Rules:
- Return only the JSON object. No markdown, no explanation.
- If the message is empty or unintelligible, set category to General and confidence to 1."""
Step 3: Build the classifier function with JSON mode
I use JSON mode to avoid regex parsing, and I set temperature low because classification is not a creative task.
import json
def classify_ticket(text: str) -> dict:
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": text},
],
response_format={"type": "json_object"},
temperature=0.1,
max_tokens=256,
)
raw = response.choices[0].message.content
return json.loads(raw)
# Test with a single example.
sample = "I was charged twice on my credit card this month. I need a refund immediately."
print(classify_ticket(sample))
Step 4: Process a batch of tickets
In production you will read from a queue or database. Here I simulate five inbound messages and print a simple table.
tickets = [
"I was charged twice on my credit card this month. I need a refund immediately.",
"How do I export my data to CSV? I cannot find the button.",
"The API returns a 500 error every time I call the embeddings endpoint.",
"Love the new UI. The dark mode is a huge improvement.",
"URGENT: our production integration is down and customers cannot check out.",
]
classified = []
for t in tickets:
try:
out = classify_ticket(t)
out["input"] = t
classified.append(out)
except Exception as e:
print(f"Failed on ticket: {e}")
# Print a simple table.
for row in classified:
print(f"{row['category']:18} | {row['priority']:8} | {row['sentiment']:8} | {row['summary']}")
Step 5: Add confidence scoring and a human-review fallback
Not every ticket should be automated. I route anything with a confidence below seven to a human review queue, and send low-priority items straight to an auto-resolve path.
def route_ticket(text: str, confidence_threshold: int = 7):
result = classify_ticket(text)
if result.get("confidence", 0) < confidence_threshold:
result["route"] = "human_review"
else:
result["route"] = "auto_resolve" if result["priority"] == "low" else "agent_queue"
return result
for t in tickets:
r = route_ticket(t, confidence_threshold=7)
print(f"{r['route']:15} | {r['category']:18} | confidence={r['confidence']} | {r['summary']}")
Run it
Running the full pipeline produces output similar to this:
agent_queue | Billing | confidence=9 | Customer requests refund for double charge agent_queue | Technical | confidence=8 | User cannot find CSV export button agent_queue | Technical | confidence=9 | API returns 500 error on embeddings endpoint auto_resolve | General | confidence=8 | Positive feedback on new dark mode UI human_review | Technical | confidence=6 | Production integration is down urgently
Wrap-up and next steps
From here, you can wire route_ticket into a webhook that listens to your support inbox, or run an evaluation pass against a few hundred hand-labeled examples to measure accuracy. If you are processing high volumes of long tickets, try deepseek-v3.2, qwen-3-32b, or kimi-k2.6 on Oxlo.ai. The flat per-request pricing keeps costs predictable even when individual messages run long. See the details at https://oxlo.ai/pricing.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.