Using LLMs for Medical Diagnosis: A Guide
We are going to build a clinical decision support agent that ingests structured patient symptoms and returns a ranked differential diagnosis with suggested workups. It targets developers prototyping triage tools or clini
We are going to build a clinical decision support agent that ingests structured patient symptoms and returns a ranked differential diagnosis with suggested workups. It targets developers prototyping triage tools or clinical training assistants. Oxlo.ai's request-based pricing keeps costs flat even when we pass in lengthy patient histories, and the OpenAI-compatible API means we can ship without rewriting any client code.
What you'll need
- Python 3.10 or newer
- The OpenAI SDK:
pip install openai - An Oxlo.ai API key from https://portal.oxlo.ai. The free tier includes enough daily requests to prototype this.
Step 1: Initialize the Oxlo.ai client
I start by instantiating the client. Because Oxlo.ai exposes a fully OpenAI-compatible endpoint, the only change from a standard OpenAI setup is the base URL.
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": "You are a helpful clinical support assistant."},
{"role": "user", "content": "Reply 'API is live' and nothing else."},
],
)
assert "live" in response.choices[0].message.content.lower()
print("Client ready")
Step 2: Write the system prompt
The prompt needs to keep the model inside a decision support scope and force valid JSON output. I also require an explicit disclaimer so the agent never presents itself as a definitive physician diagnosis.
SYSTEM_PROMPT = """You are a clinical decision support agent. Your job is to read a structured patient intake and return a JSON object with exactly these keys:
- 'differential': a list of the 3 most likely conditions ranked by likelihood. Each item must have 'condition' and 'rationale'.
- 'red_flags': a list of any urgent symptoms or findings that require immediate in-person evaluation.
- 'next_steps': a list of recommended labs, imaging, or referrals.
Rules:
1. Always include a disclaimer that this is not a medical diagnosis.
2. If the presentation is vague, state the uncertainty clearly.
3. Return only valid JSON. Do not wrap the output in markdown fences."""
Step 3: Format the intake and call the model
I serialize the patient data to JSON and call Llama 3.3 70B with JSON mode enabled. Oxlo.ai supports this on compatible models, so the output is guaranteed to be parseable. If you need vision later, Kimi K2.6 handles image inputs on the same endpoint.
import json
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
def build_patient_message(age, sex, chief_complaint, history, vitals):
payload = {
"patient_profile": {"age": age, "sex": sex, "vitals": vitals},
"chief_complaint": chief_complaint,
"history": history
}
return json.dumps(payload, indent=2)
def get_differential(patient_json: str):
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": patient_json},
],
response_format={"type": "json_object"},
temperature=0.2,
)
return json.loads(response.choices[0].message.content)
patient_input = build_patient_message(
age=34,
sex="female",
chief_complaint="Sharp right lower quadrant pain for 6 hours, nausea, no fever",
history="No prior surgeries. Last menstrual period 2 weeks ago. No known drug allergies.",
vitals={"bp": "118/76", "hr": 96, "temp_c": 37.1, "rr": 16}
)
result = get_differential(patient_input)
print(json.dumps(result, indent=2))
Step 4: Validate and harden the agent
Before showing anything to a user, I check that the required keys exist and inject a fallback disclaimer if the model forgets it. This keeps the interface predictable.
REQUIRED_KEYS = {"differential", "red_flags", "next_steps"}
def safe_diagnose(patient_json: str):
try:
data = get_differential(patient_json)
except Exception as e:
return {"error": f"Parsing failed: {e}", "raw": None}
missing = REQUIRED_KEYS - set(data.keys())
if missing:
return {"error": f"Missing keys: {missing}", "raw": data}
text_blob = json.dumps(data).lower()
if "not a medical diagnosis" not in text_blob and "not a substitute" not in text_blob:
data["disclaimer"] = "This output is not a substitute for professional medical advice."
return data
final = safe_diagnose(patient_input)
print(json.dumps(final, indent=2))
Run it
Here is the full script. Execute it with your Oxlo.ai key exported and you should see a structured differential for the sample case.
import json
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
SYSTEM_PROMPT = """You are a clinical decision support agent. Your job is to read a structured patient intake and return a JSON object with exactly these keys:
- 'differential': a list of the 3 most likely conditions ranked by likelihood. Each item must have 'condition' and 'rationale'.
- 'red_flags': a list of any urgent symptoms or findings that require immediate in-person evaluation.
- 'next_steps': a list of recommended labs, imaging, or referrals.
Rules:
1. Always include a disclaimer that this is not a medical diagnosis.
2. If the presentation is vague, state the uncertainty clearly.
3. Return only valid JSON. Do not wrap the output in markdown fences."""
def build_patient_message(age, sex, chief_complaint, history, vitals):
payload = {
"patient_profile": {"age": age, "sex": sex, "vitals": vitals},
"chief_complaint": chief_complaint,
"history": history
}
return json.dumps(payload, indent=2)
def get_differential(patient_json: str):
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": patient_json},
],
response_format={"type": "json_object"},
temperature=0.2,
)
return json.loads(response.choices[0].message.content)
REQUIRED_KEYS = {"differential", "red_flags", "next_steps"}
def safe_diagnose(patient_json: str):
try:
data = get_differential(patient_json)
except Exception as e:
return {"error": f"Parsing failed: {e}", "raw": None}
missing = REQUIRED_KEYS - set(data.keys())
if missing:
return {"error": f"Missing keys: {missing}", "raw": data}
text_blob = json.dumps(data).lower()
if "not a medical diagnosis" not in text_blob and "not a substitute" not in text_blob:
data["disclaimer"] = "This output is not a substitute for professional medical advice."
return data
if __name__ == "__main__":
patient_input = build_patient_message(
age=34,
sex="female",
chief_complaint="Sharp right lower quadrant pain for 6 hours, nausea, no fever",
history="No prior surgeries. Last menstrual period 2 weeks ago. No known drug allergies.",
vitals={"bp": "118/76", "hr": 96, "temp_c": 37.1, "rr": 16}
)
out = safe_diagnose(patient_input)
print(json.dumps(out, indent=2))
Example output:
{
"differential": [
{
"condition": "Appendicitis",
"rationale": "Classic presentation of acute RLQ pain with nausea in a young adult."
},
{
"condition": "Ruptured ovarian cyst",
"rationale": "Fits demographic and timing with recent LMP; absence of fever makes infection less likely."
},
{
"condition": "Ectopic pregnancy",
"rationale": "Must be ruled out in any female of reproductive age with acute abdominal pain."
}
],
"red_flags": [
"Possible ectopic pregnancy requires immediate beta-hCG and pelvic ultrasound.",
"Peritoneal signs or worsening pain warrant emergency department evaluation now."
],
"next_steps": [
"CBC, CMP, urinalysis",
"Quantitative beta-hCG",
"Transvaginal pelvic ultrasound",
"Surgical consult if imaging confirms appendicitis"
],
"disclaimer": "This output is not a substitute for professional medical advice."
}
Next steps
Wire this agent into a FastAPI endpoint so a frontend intake form can POST directly to Oxlo.ai, or add a retrieval layer over a medical knowledge base to ground each rationale in current literature. If you start processing long electronic health record extracts, keep an eye on Oxlo.ai's pricing; the flat per-request cost avoids the token bill surprises common with long-context workloads.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.