A Practical Guide to Using LLMs in Social Science
Social scientists often spend weeks manually coding interview transcripts. In this guide, I will walk you through building a thematic analysis agent that reads transcripts, generates inductive codes, and synthesizes them
Social scientists often spend weeks manually coding interview transcripts. In this guide, I will walk you through building a thematic analysis agent that reads transcripts, generates inductive codes, and synthesizes them into a structured report. The pipeline uses Oxlo.ai and its request-based pricing, so processing long participant narratives does not inflate your inference costs.
What you'll need
You will need Python 3.10 or newer, the OpenAI SDK, and an Oxlo.ai API key.
pip install openai
Grab your API key from https://portal.oxlo.ai. We will use it to call Oxlo.ai's OpenAI-compatible endpoints.
Step 1: Set up the client and load a transcript
I start by initializing the Oxlo.ai client and loading a short interview excerpt. The agent works on full transcripts because Oxlo.ai charges per request, not per token.
from openai import OpenAI
client = OpenAI(base_url="https://api.oxlo.ai/v1", api_key="YOUR_OXLO_API_KEY")
TRANSCRIPT = """
Interviewer: Can you describe how remote work has changed your daily routine?
Participant: Well, I used to have a clear boundary. I would leave the house at eight and come back at six. Now, my laptop is always there. I find myself checking emails at ten PM. My partner says I am never fully present. But I also save two hours on commuting, which I try to spend with my kids.
Interviewer: How do you manage that tension?
Participant: It is hard. I feel guilty when I close the laptop at five because my colleagues are still online. There is this unspoken pressure to be available. On the other hand, my kids are happier when I am home. It is a trade-off I have not figured out yet.
"""
Step 2: Chunk the transcript into episodes
Long interviews are easier to code in small, topical chunks. This function splits the transcript on interviewer questions and tags each segment.
import re
def chunk_transcript(text):
parts = re.split(r"(?=Interviewer:)", text.strip())
episodes = []
for i, part in enumerate(parts):
if not part.strip():
continue
episodes.append({"id": f"ep_{i}", "text": part.strip()})
return episodes
episodes = chunk_transcript(TRANSCRIPT)
print(f"Loaded {len(episodes)} episodes")
Step 3: Define the coding agent system prompt
The system prompt tells the model to act as a qualitative researcher performing inductive open coding. It must return structured JSON.
SYSTEM_PROMPT = '''You are a qualitative research assistant performing inductive thematic coding on interview transcripts.
Instructions:
1. Read the provided interview episode carefully.
2. Identify 2 to 5 distinct concepts, feelings, or behaviors mentioned by the participant.
3. For each concept, provide:
- code: a short label in snake_case (max 3 words)
- definition: one sentence explaining what this code captures
- evidence: the exact quote from the text that supports the code
4. Return ONLY a valid JSON object with a single key "codes" containing a list of code objects. Do not wrap the JSON in markdown fences.
Example format:
{
"codes": [
{"code": "boundary_blur", "definition": "Participant describes fading separation between work and home life.", "evidence": "my laptop is always there"}
]
}
'''
Step 4: Run first-pass coding on every episode
Now I send each episode to Qwen 3 32B on Oxlo.ai. It handles structured JSON and reasoning tasks reliably. The function extracts the JSON and attaches metadata.
import json
def code_episode(episode):
user_message = f"Episode ID: {episode['id']}\n\nText:\n{episode['text']}"
response = client.chat.completions.create(
model="qwen-3-32b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": user_message},
],
temperature=0.2,
)
raw = response.choices[0].message.content
cleaned = raw.strip().removeprefix("
```json").removeprefix("```
").removesuffix("
```").strip()
return json.loads(cleaned)
coded_episodes = []
for ep in episodes:
try:
result = code_episode(ep)
coded_episodes.append({"id": ep["id"], "codes": result.get("codes", [])})
print(f"Coded {ep['id']}: {len(result.get('codes', []))} codes")
except Exception as e:
print(f"Failed on {ep['id']}: {e}")
Step 5: Synthesize codes into cross-cutting themes
After coding, we reduce the code list into higher-order themes. I pass the aggregated codes to Llama 3.3 70B for general-purpose synthesis.
def synthesize_themes(coded_eps):
lines = []
for ce in coded_eps:
for c in ce["codes"]:
lines.append(f"- {c['code']}: {c['definition']} (from {ce['id']})")
aggregation = "\n".join(lines)
prompt = f"""You are a senior qualitative researcher. Given the following inductively generated codes from multiple interview episodes, synthesize them into 3 to 6 overarching themes.
For each theme, provide:
- theme: a concise title
- description: one paragraph explaining the theme
- related_codes: list of codes that belong under this theme
- key_quotes: up to 2 representative quotes from the original evidence
Return ONLY valid JSON with a single key "themes" containing a list of theme objects.
Codes:
{aggregation}
"""
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": "You synthesize qualitative data into structured themes. Always return valid JSON."},
{"role": "user", "content": prompt},
],
temperature=0.3,
)
raw = response.choices[0].message.content.strip()
cleaned = raw.removeprefix("```
json").removeprefix("
```").removesuffix("```
").strip()
return json.loads(cleaned)
themes = synthesize_themes(coded_episodes)
print(json.dumps(themes, indent=2))
Step 6: Generate the final report
Finally, I render a human-readable markdown report from the structured themes and save it to disk.
def write_report(themes, filename="thematic_report.md"):
with open(filename, "w") as f:
f.write("# Thematic Analysis Report\n\n")
for i, t in enumerate(themes.get("themes", []), 1):
f.write(f"## Theme {i}: {t['theme']}\n\n")
f.write(f"**Description:** {t['description']}\n\n")
f.write("**Related codes:** " + ", ".join(t.get("related_codes", [])) + "\n\n")
f.write("**Key quotes:**\n")
for q in t.get("key_quotes", []):
f.write(f"- \"{q}\"\n")
f.write("\n")
print(f"Report written to {filename}")
write_report(themes)
Run it
Putting it all together, the main block below chunks the transcript, codes each episode, synthesizes themes, and writes the report.
if __name__ == "__main__":
episodes = chunk_transcript(TRANSCRIPT)
print(f"Loaded {len(episodes)} episodes")
coded = []
for ep in episodes:
result = code_episode(ep)
coded.append({"id": ep["id"], "codes": result.get("codes", [])})
print(f"Coded {ep['id']}")
themes = synthesize_themes(coded)
write_report(themes)
for t in themes.get("themes", []):
print(f"Theme: {t['theme']} | Codes: {', '.join(t.get('related_codes', []))}")
Expected output:
Loaded 2 episodes Coded ep_0: 3 codes Coded ep_1: 3 codes Report written to thematic_report.md Theme: Boundary Blur and Availability Pressure | Codes: boundary_blur, availability_pressure Theme: Guilt and Trade-offs | Codes: guilt, work_life_tension Theme: Commute Redistribution | Codes: commute_redistribution, family_presence
Wrap-up and next steps
You now have a working thematic analysis agent that turns raw transcripts into structured reports. Two concrete ways to extend it: first, add deductive coding by loading an existing codebook and asking the model to map episodes to prior codes alongside inductive ones. Second, batch-process an entire corpus by reading .txt files from a directory and running the pipeline with a ThreadPoolExecutor to exploit Oxlo.ai's no-cold-start endpoints.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.