Dev.to AI 🤖 Ai 👁 0 📖 2 min read

Using LLMs for Image Classification: Techniques and Applications

Image classification has traditionally relied on convolutional neural networks trained on fixed label sets. Today, vision-capable large language models let you classify images through natural language instructions, elimi

Image classification has traditionally relied on convolutional neural networks trained on fixed label sets. Today, vision-capable large language models let you classify images through natural language instructions, eliminating the need for task-specific training data in many cases. These models accept an image and a text prompt, then return a label, description, or structured analysis. For developers running classification at scale, the inference platform matters as much as the model. Oxlo.ai provides vision models, request-based pricing, and full OpenAI SDK compatibility, making it straightforward to build production image classification pipelines without token-cost surprises.

The Architecture Behind Vision Classification

Modern vision-language models combine a vision encoder with a large language model backbone. The encoder processes an image into a sequence of patch embeddings, which a projection layer aligns to the language model's token embedding space. From there, the transformer decoder processes text and image tokens through shared attention layers. Models available on Oxlo.ai, such as Gemma 3 27B and Kimi VL A3B, handle high-resolution inputs and generate text outputs conditioned on both modalities. Because the entire forward pass happens in a single inference request, platform latency and pricing structure directly affect throughput.

Zero-Shot Classification

Zero-shot classification is the simplest way to deploy an LLM for image classification. You provide a natural language description of candidate labels, and the model selects the best fit without any gradient updates. This works because vision-language models are trained on broad image-text corpora that embed semantic relationships between visual concepts and language.

Below is a complete example using the Oxlo.ai API with the OpenAI Python SDK. Replace the model identifier with the exact version from your Oxlo.ai dashboard if needed.

import os
from openai import OpenAI

client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.environ["OXLO_API_KEY"]
)

response = client.chat.completions.create(
model="gemma-3-27b",
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": "Classify this image into exactly one category: electronics, clothing, or food. Respond with only the category name."
},
{
"type": "image_url",
"image_url": {"url": "https://example.com/product-image.jpg"}
}
]
}
],
max_tokens=20,
temperature

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.