Dev.to AI 🤖 Ai 👁 0 📖 4 min read

LLM Models for Multimodal Tasks

In this tutorial we will build a multimodal bug triage agent that accepts a screenshot and a short description, then returns structured JSON with severity, component, and reproduction steps. It runs entirely on Oxlo.ai's

In this tutorial we will build a multimodal bug triage agent that accepts a screenshot and a short description, then returns structured JSON with severity, component, and reproduction steps. It runs entirely on Oxlo.ai's vision-capable models under flat per-request pricing, so sending a 4K screenshot costs the same as a one-line text prompt. If your team currently routes bug reports by hand, this agent eliminates that bottleneck.

What you'll need

  • Python 3.10 or newer.
  • The OpenAI SDK installed with pip install openai.
  • An Oxlo.ai API key from https://portal.oxlo.ai. The free tier includes 60 requests per day and a 7-day full-access trial, which is enough to prototype this agent. See https://oxlo.ai/pricing for plan details.
  • A sample screenshot saved as bug_screenshot.png in your working directory.

Step 1: Verify the text baseline with Llama 3.3 70B

Before adding images, I verify that the Oxlo.ai endpoint and my API key are working with a simple text-only request. I use Llama 3.3 70B because it is a reliable general-purpose model for quick sanity checks.

from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key=os.environ.get("OXLO_API_KEY")
)

response = client.chat.completions.create(
    model="llama-3.3-70b",
    messages=[
        {"role": "user", "content": "Hello, confirm you are working."}
    ]
)

print(response.choices[0].message.content)

Step 2: Add vision input with base64 encoding

I load a local screenshot, encode it to base64, and send it to Kimi K2.6, which handles both image understanding and long-context reasoning. Oxlo.ai serves this with no cold starts, so the first request after idle time returns immediately.

import base64
from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key=os.environ.get("OXLO_API_KEY")
)

def encode_image(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

base64_image = encode_image("bug_screenshot.png")

response = client.chat.completions.create(
    model="kimi-k2.6",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe what you see in this screenshot."},
                {
                    "type": "image_url",
                    "image_url": {
                        "url": f"data:image/png;base64,{base64_image}"
                    }
                }
            ]
        }
    ]
)

print(response.choices[0].message.content)

Step 3: Write the system prompt

The system prompt constrains the model to output only the fields we need and explicitly forbids hallucination. Keeping this prompt tight reduces token usage and improves consistency.

SYSTEM_PROMPT = """You are a bug triage agent. You analyze software bug reports that contain a screenshot and a short description.

Output a JSON object with these exact fields:
- summary: one sentence describing the bug
- severity: one of critical, high, medium, low
- component: the likely UI or backend component affected
- reproduction_steps: a numbered list of steps inferred from the screenshot and description
- question_for_reporter: one clarifying question if information is missing, otherwise null

Rules:
- Be concise.
- Do not hallucinate details not supported by the input.
- Output only the JSON object."""

Step 4: Enforce structured output with JSON mode

I wrap the text and vision inputs into a single function and set response_format to json_object. Because Oxlo.ai charges a flat rate per request, the cost stays the same whether the screenshot is 100 KB or 4 MB, which makes this workflow predictable at scale.

import json
import base64
from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key=os.environ.get("OXLO_API_KEY")
)

SYSTEM_PROMPT = """You are a bug triage agent. You analyze software bug reports that contain a screenshot and a short description.

Output a JSON object with these exact fields:
- summary: one sentence describing the bug
- severity: one of critical, high, medium, low
- component: the likely UI or backend component affected
- reproduction_steps: a numbered list of steps inferred from the screenshot and description
- question_for_reporter: one clarifying question if information is missing, otherwise null

Rules:
- Be concise.
- Do not hallucinate details not supported by the input.
- Output only the JSON object."""

def encode_image(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

def analyze_bug(image_path: str, description: str) -> dict:
    base64_image = encode_image(image_path)

    response = client.chat.completions.create(
        model="kimi-k2.6",
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            {
                "role": "user",
                "content": [
                    {"type": "text", "text": f"Bug description: {description}"},
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/png;base64,{base64_image}"
                        }
                    }
                ]
            }
        ],
        response_format={"type": "json_object"}
    )

    return json.loads(response.choices[0].message.content)

Run it

I call analyze_bug with a sample screenshot and a short description. The agent returns structured data in a few seconds.

result = analyze_bug(
    "bug_screenshot.png",
    "The checkout button does nothing after I fill out the shipping form."
)

print(json.dumps(result, indent=2))

Example output:

{
  "summary": "Checkout button becomes unresponsive after shipping form completion",
  "severity": "high",
  "component": "checkout-ui",
  "reproduction_steps": [
    "Navigate to the checkout page",
    "Fill out the shipping form with valid details",
    "Click the 'Complete Purchase' button"
  ],
  "question_for_reporter": "Does the browser console show any JavaScript errors when the button is clicked?"
}

Wrap-up

You now have a working multimodal agent that turns unstructured bug reports into structured tickets. Next, you can wire this function into a Slack bot or GitHub webhook so it runs automatically when a user uploads a screenshot. If you need stronger multilingual reasoning for international reports, swap Kimi K2.6 for Qwen 3 32B with no other code changes.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.