Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Unlocking LLM Potential in Social Science Research

Social scientists can spend weeks manually coding interview transcripts to extract themes. In this tutorial, I will build a qualitative coding assistant that ingests a full transcript and returns structured themes, suppo

Social scientists can spend weeks manually coding interview transcripts to extract themes. In this tutorial, I will build a qualitative coding assistant that ingests a full transcript and returns structured themes, supporting quotes, and an analytic memo in one pass. I run it on Oxlo.ai, where flat per-request pricing makes it practical to feed in long transcripts without tracking token costs.

What you'll need

Step 1: Configure the Oxlo.ai client

I point the OpenAI SDK at Oxlo.ai and select Kimi K2.6. Its 131K context window handles full interview transcripts, and the flat per-request rate means a forty page transcript costs the same as a one sentence prompt.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key="YOUR_OXLO_API_KEY"
)

MODEL = "kimi-k2.6"

Step 2: Lock down the system prompt

The system prompt is the only part a researcher needs to edit. I instruct the model to behave as a thematic analyst, map every theme to exact quotes, and return strict JSON.

SYSTEM_PROMPT = """You are a qualitative research assistant specializing in thematic analysis of interview transcripts.

Follow these rules:
1. Read the entire transcript carefully.
2. Identify 3 to 5 distinct themes relevant to the participant's experience.
3. For each theme, provide:
   - theme_name: a concise label
   - description: 1 to 2 sentences explaining the theme
   - evidence: an array of exact quotes from the transcript that support the theme
4. Write a 100 word analytic memo summarizing how the themes relate to one another.
5. Return ONLY a JSON object with keys: themes (array), memo (string).

Do not paraphrase quotes. Use the speaker's exact words."""

Step 3: Build the analysis function

I wrap the API call so it accepts raw transcript text and returns parsed JSON. I enable JSON mode to avoid regex cleanup.

import json

def code_transcript(transcript: str):
    response = client.chat.completions.create(
        model=MODEL,
        response_format={"type": "json_object"},
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": transcript},
        ],
    )
    raw = response.choices[0].message.content
    return json.loads(raw)

Step 4: Feed it a sample interview

Here is a synthetic transcript about remote work. I pass the raw text straight to the function.

TRANSCRIPT = """
Interviewer: How has working from home affected your daily routine?
Participant: Honestly, I lost all boundaries. My kitchen became my office, and I found myself answering emails at 10 PM. 
Interviewer: Has that changed over time?
Participant: Yes. I started blocking my calendar for lunch walks. That small ritual brought back a sense of control. 
Interviewer: Any impact on collaboration?
Participant: Video calls feel exhausting. I miss spontaneous hallway conversations. But async documentation has actually made our team more thoughtful.
Interviewer: Would you return to the office full time?
Participant: Only if it is hybrid. I need the flexibility, but I also miss the energy of being around people.
"""

result = code_transcript(TRANSCRIPT)
print(json.dumps(result, indent=2))

Run it

Executing the script produces structured output like this. Every theme anchors to an exact quote, and the memo connects boundary loss with the evolution toward hybrid preferences.

{
  "themes": [
    {
      "theme_name": "Erosion of Work-Life Boundaries",
      "description": "The participant describes how remote work dissolved physical and temporal boundaries between personal and professional life.",
      "evidence": [
        "I lost all boundaries. My kitchen became my office, and I found myself answering emails at 10 PM."
      ]
    },
    {
      "theme_name": "Intentional Rituals for Autonomy",
      "description": "The participant developed deliberate practices to reclaim control over their schedule and wellbeing.",
      "evidence": [
        "I started blocking my calendar for lunch walks. That small ritual brought back a sense of control."
      ]
    },
    {
      "theme_name": "Ambivalence Toward Digital Collaboration",
      "description": "The participant experiences video fatigue but acknowledges improved thoughtfulness through asynchronous communication.",
      "evidence": [
        "Video calls feel exhausting. I miss spontaneous hallway conversations. But async documentation has actually made our team more thoughtful."
      ]
    },
    {
      "theme_name": "Hybrid as Optimal Compromise",
      "description": "The participant explicitly prefers a mixed model that preserves flexibility while satisfying a need for social energy.",
      "evidence": [
        "Only if it is hybrid. I need the flexibility, but I also miss the energy of being around people."
      ]
    }
  ],
  "memo": "The participant's experience traces an arc from boundary loss to deliberate recovery, culminating in a negotiated preference for hybrid work. Themes of autonomy and ambivalence intersect around control: digital tools both enable and constrain collaboration, while physical space serves as a symbolic marker separating work from life. The desire for hybrid arrangements reflects not nostalgia for the office, but a calibrated strategy to optimize flexibility and social connection."
}

Wrap-up

This agent replaces hours of manual highlighting with a single API call. Two concrete next steps: loop over a directory of transcripts to batch code an entire study, or tighten the system prompt to enforce a specific theoretical framework such as constructivist grounded theory. For projects that process hundreds of long interviews, Oxlo.ai's per-request pricing keeps costs predictable as volume scales.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.