Mitigating Bias in Large Language Models
We are building a lightweight bias audit agent that scores text for demographic bias and returns a neutral rewrite. It is useful for teams shipping LLM features who need a fast, structured guardrail before content reache
We are building a lightweight bias audit agent that scores text for demographic bias and returns a neutral rewrite. It is useful for teams shipping LLM features who need a fast, structured guardrail before content reaches users. Because Oxlo.ai uses flat per-request pricing rather than per token, detailed at https://oxlo.ai/pricing, and serves popular models with no cold starts, running this on long transcripts or documents stays fast and cost-predictable.
What you'll need
- An Oxlo.ai API key from https://portal.oxlo.ai
- Python 3.10 or newer
- The OpenAI SDK installed with
pip install openai
Step 1: Scaffold the Oxlo.ai client
I start by verifying connectivity to Oxlo.ai and confirming the API key is active. A quick ping against Llama 3.3 70B proves the client is ready.
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Say 'Connection OK' and nothing else."},
],
)
print(response.choices[0].message.content)
Step 2: Write the system prompt
The system prompt turns the model into a strict audit API. It must analyze the input for demographic bias and return only a JSON object containing a verdict, severity score, explanation, and neutral rewrite.
SYSTEM_PROMPT = """You are a bias audit engine.
Analyze the user text for demographic bias including gender, age, ethnicity, religion, disability, or socioeconomic status.
Return ONLY a JSON object with this exact shape:
{
"biased": boolean,
"severity": integer from 1 to 5,
"category": "none" or the type of bias detected,
"explanation": "detailed reasoning",
"rewrite": "neutral version if biased, else original text"
}
Do not include markdown fences or commentary outside the JSON."""
Step 3: Build the audit function
I wrap the call in a small Python function that injects the system prompt, sends the text to Oxlo.ai, and parses the JSON response. Keeping this in a pure function makes it easy to drop into a FastAPI route or Celery task later.
import json
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
SYSTEM_PROMPT = """You are a bias audit engine.
Analyze the user text for demographic bias including gender, age, ethnicity, religion, disability, or socioeconomic status.
Return ONLY a JSON object with this exact shape:
{
"biased": boolean,
"severity": integer from 1 to 5,
"category": "none" or the type of bias detected,
"explanation": "detailed reasoning",
"rewrite": "neutral version if biased, else original text"
}
Do not include markdown fences or commentary outside the JSON."""
def audit_text(text: str) -> dict:
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": text},
],
)
raw = response.choices[0].message.content.strip()
if raw.startswith("
```"):
raw = raw.split("\n", 1)[1].rsplit("```
", 1)[0].strip()
return json.loads(raw)
# Quick sanity check
sample = "The chairman should discuss this with his team."
print(audit_text(sample))
Step 4: Run a batch evaluation
To see how the agent handles edge cases, I pass a list of sentences ranging from overtly biased to neutral. Because Oxlo.ai uses flat per-request pricing, long inputs do not increase cost, so I can feed full paragraphs without worrying about token count.
import json
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
SYSTEM_PROMPT = """You are a bias audit engine.
Analyze the user text for demographic bias including gender, age, ethnicity, religion, disability, or socioeconomic status.
Return ONLY a JSON object with this exact shape:
{
"biased": boolean,
"severity": integer from 1 to 5,
"category": "none" or the type of bias detected,
"explanation": "detailed reasoning",
"rewrite": "neutral version if biased, else original text"
}
Do not include markdown fences or commentary outside the JSON."""
def audit_text(text: str) -> dict:
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": text},
],
)
raw = response.choices[0].message.content.strip()
if raw.startswith("
```"):
raw = raw.split("\n", 1)[1].rsplit("```
", 1)[0].strip()
return json.loads(raw)
candidates = [
"The chairman should discuss this with his team.",
"Young people are too lazy to learn new tools.",
"Please send the meeting notes to all attendees after the call.",
"The disabled employee needs special help to keep up.",
"We are looking for a native English speaker to sound professional.",
]
for text in candidates:
result = audit_text(text)
print(f"Text: {text}")
print(f"Verdict: {result}")
print()
Run it
Save the batch script as audit.py, export your key, and run python audit.py. You should see structured JSON verdicts similar to these:
Text: The chairman should discuss this with his team.
Verdict: {'biased': True, 'severity': 2, 'category': 'gender', 'explanation': 'Uses masculine nouns and pronouns assuming male leadership.', 'rewrite': 'The chair should discuss this with their team.'}
Text: Young people are too lazy to learn new tools.
Verdict: {'biased': True, 'severity': 3, 'category': 'age', 'explanation': 'Generalizes a negative trait across an entire age group.', 'rewrite': 'Some individuals may need additional motivation to learn new tools.'}
Text: Please send the meeting notes to all attendees after the call.
Verdict: {'biased': False, 'severity': 1, 'category': 'none', 'explanation': 'No demographic group is targeted or stereotyped.', 'rewrite': 'Please send the meeting notes to all attendees after the call.'}
Next steps
Wrap the audit_text function in a FastAPI middleware so every LLM response in your stack gets screened before it reaches the user. Then build a small labeled evaluation set and compare Llama 3.3 70B against Qwen 3 32B or DeepSeek V3.2 on Oxlo.ai to see which model gives the best precision-recall trade-off for your domain.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.