Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 4 min read

Integrating LLM with Web Applications: A Beginner's Guide

Integrating a large language model into a web application is no longer limited to AI-native products. Every developer-facing tool, support portal, and content platform can benefit from contextual generation, but the inte

Integrating a large language model into a web application is no longer limited to AI-native products. Every developer-facing tool, support portal, and content platform can benefit from contextual generation, but the integration path is often obscured by provider-specific SDKs, unpredictable token costs, and cold-start latency. A robust integration starts with treating the LLM as an external inference service, consumed through a clean backend layer that handles authentication, retries, and response streaming.

Architecture Overview

A typical full-stack LLM integration follows a three-tier pattern. The client browser communicates with your application server, which in turn calls an inference API. This indirection keeps API keys out of the frontend, lets you implement caching and rate limiting, and allows you to swap model providers without touching client code. For the inference layer, you need an endpoint that speaks a standard format. Oxlo.ai exposes an OpenAI-compatible API at https://api.oxlo.ai/v1, which means your backend can use the official OpenAI SDKs while routing requests to Oxlo.ai's infrastructure.

Choosing an Inference Provider

When selecting a backend, evaluate model variety, latency, and pricing mechanics. Token-based billing can make long-context chat sessions and agentic loops prohibitively expensive because every input token counts against your budget. Oxlo.ai uses request-based pricing with one flat cost per API call regardless of prompt length. For web applications that send large HTML chunks, conversation history, or retrieved documents with every request, this model can reduce costs significantly compared to token-based alternatives. Oxlo.ai also offers 45+ open-source and proprietary models, from general-purpose options like Llama 3.3 70B and Qwen 3 32B to specialized coders and vision models, with no cold starts on popular deployments.

Backend Implementation

Here is a minimal Node.js/Express backend that proxies chat requests to Oxlo.ai. The implementation uses the OpenAI Node.js SDK, which works because Oxlo.ai's API is fully compatible.

import express from 'express';
import OpenAI from 'openai';

const app = express();
app.use(express.json());

const client = new OpenAI({
  apiKey: process.env.OXLO_API_KEY,
  baseURL: 'https://api.oxlo.ai/v1',
});

app.post('/api/chat', async (req, res) => {
  const { messages, model = 'llama-3.3-70b' } = req.body;

  try {
    const stream = await client.chat.completions.create({
      model,
      messages,
      stream: true,
    });

    res.setHeader('Content-Type', 'text/event-stream');
    res.setHeader('Cache-Control', 'no-cache');
    res.setHeader('Connection', 'keep-alive');

    for await (const chunk of stream) {
      const content = chunk.choices[0]?.delta?.content || '';
      res.write(`data: ${JSON.stringify({ content })}\n\n`);
    }

    res.write('data: [DONE]\n\n');
    res.end();
  } catch (error) {
    res.status(500).json({ error: error.message });
  }
});

app.listen(3000);

This pattern keeps your Oxlo.ai key on the server and streams tokens to the client as they are generated.

Frontend Integration

On the frontend, you consume the stream using the EventSource API or a fetch-based reader. Below is a React hook pattern that appends chunks to a message state.

import { useState, useCallback } from 'react';

export function useChat() {
  const [messages, setMessages] = useState([]);

  const sendMessage = useCallback(async (text) => {
    const userMsg = { role: 'user', content: text };
    setMessages((prev) => [...prev, userMsg]);

    const res = await fetch('/api/chat', {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({ messages: [...messages, userMsg] }),
    });

    const reader = res.body.getReader();
    const decoder = new TextDecoder();
    let assistantContent = '';

    setMessages((prev) => [...prev, { role: 'assistant', content: '' }]);

    while (true) {
      const { done, value } = await reader.read();
      if (done) break;

      const chunk = decoder.decode(value, { stream: true });
      const lines = chunk.split('\n').filter((line) => line.startsWith('data: '));

      for (const line of lines) {
        const data = line.replace('data: ', '');
        if (data === '[DONE]') continue;
        const parsed = JSON.parse(data);
        assistantContent += parsed.content;

        setMessages((prev) => {
          const updated = [...prev];
          updated[updated.length - 1] = {
            role: 'assistant',
            content: assistantContent,
          };
          return updated;
        });
      }
    }
  }, [messages]);

  return { messages, sendMessage };
}

Advanced Patterns: Tool Use and Structured Output

Modern web applications need more than plain text generation. Function calling lets the model request external data, and JSON mode forces structured output for forms and dashboards. Oxlo.ai supports both capabilities across compatible models. You can define a tool schema in your backend request and handle the loop server-side before returning a final rendered result to the browser.

For applications that process screenshots or diagrams, vision models such as Gemma 3 27B and Kimi VL A3B on Oxlo.ai accept base64-encoded images through the same chat completions endpoint. Because Oxlo.ai charges per request rather than per image token, multimodal workloads with high-resolution inputs remain predictable.

Cost and Scaling Considerations

The primary cost driver in most LLM integrations is not the model choice but the pricing mechanics. Token-based providers scale costs linearly with input length, which penalizes applications that pass long documents, thread history, or retrieved context into every prompt. Oxlo.ai's flat per-request pricing removes this variable. For a web app with growing context windows or agentic step-by-step workflows, this can make operational costs significantly more predictable. You can start building on the free tier, which includes 60 requests per day and access to more than 16 models, then scale through Pro or Premium plans as traffic grows. See https://oxlo.ai/pricing for current plan details.

Conclusion

Integrating an LLM into your web stack is fundamentally an API integration problem. By placing an OpenAI-compatible provider like Oxlo.ai behind your own backend, you retain full control over authentication, caching, and user experience while gaining access to a broad model catalog under a request-based pricing model that favors long-context web applications. Start with a streaming chat endpoint, add structured output or vision when needed, and let your backend abstraction shield the frontend from provider details.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.