Dev.to AI 🤖 Ai 👁 0 📖 4 min read

Deploying LLM Models on Cloud Platforms

Deploying large language models in production requires more than downloading weights from Hugging Face. Engineering teams must navigate GPU driver compatibility, tensor parallelism configuration, auto-scaling policies, a

Deploying large language models in production requires more than downloading weights from Hugging Face. Engineering teams must navigate GPU driver compatibility, tensor parallelism configuration, auto-scaling policies, and continuous batching parameters before the first token reaches a client. Cloud platforms like AWS, Google Cloud, and Azure provide the raw compute, but the operational burden of self-hosting remains significant. This article examines the practical paths to production inference, from raw VM deployment to managed APIs, and explains where each approach fits.

Self-Hosted Deployment on Cloud VMs

The most flexible approach is provisioning dedicated GPU instances and serving models with an open inference engine such as vLLM, TGI, or TensorRT-LLM. On AWS, this typically begins with a g6e or p5 instance, while GCP offers A3 VMs and Azure provides NC-series virtual machines. After provisioning, you install NVIDIA drivers, CUDA, and the serving framework, then download and shard model weights across GPUs.

A minimal vLLM deployment on a single node looks like this:

docker run --runtime nvidia --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.3-70B-Instruct \
  --tensor-parallel-size 8 \
  --max-model-len 8192

This gives you full control over quantization, speculative decoding, and custom scheduling. However, it also means you own the lifecycle: OS patching, driver upgrades, spot instance interruption handling, and scaling logic. For long-context workloads, memory pressure spikes unpredictably, often forcing over-provisioning that drives up cost.

Kubernetes and Orchestration Layers

Teams running multiple models usually graduate to Kubernetes with GPU operators and custom schedulers. You define InferenceService CRDs with KServe or deploy Ray Serve clusters to handle model composition and multi-node pipelines. The configuration complexity increases sharply.

A simplified KServe manifest might look like this:

apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: llama-3-3-70b
spec:
  predictor:
    nodeSelector:
      node.kubernetes.io/instance-type: p5.48xlarge
    containers:
      - name: kserve-container
        image: vllm/vllm-openai:latest
        args:
          - --model
          - meta-llama/Llama-3.3-70B-Instruct
          - --tensor-parallel-size
          - "8"

While Kubernetes enables replication and rolling updates, it introduces new failure modes. GPU health checks, pod preemption, and HPA thresholds based on GPU utilization rather than request queue depth are common sources of production incidents. Cold starts remain a problem if you scale to zero, yet keeping replicas warm burns budget during low traffic.

Managed API Inference with Oxlo.ai

For many production workloads, the operational overhead of self-hosting outweighs the benefits of direct hardware control. Managed inference APIs abstract away the cluster, the drivers, and the scheduling logic, letting engineers focus on prompt engineering, evaluation, and product integration.

Oxlo.ai offers a developer-first alternative built on request-based pricing. Instead of metering tokens, Oxlo.ai charges one flat cost per API request regardless of prompt length. This model eliminates the cost unpredictability that comes with long-context and agentic workloads, where input tokens can dominate the bill on token-based platforms. You can explore the exact structure on the Oxlo.ai pricing page.

The platform hosts more than 45 open-source and proprietary models across seven categories, including general-purpose LLMs like Llama 3.3 70B and Qwen 3 32B, reasoning models such as DeepSeek R1 671B MoE and Kimi K2.6, and specialized endpoints for code, vision, audio, and embeddings. Because Oxlo.ai is fully compatible with the OpenAI SDK, migration from another provider or from a local OpenAI-compatible server requires only a base URL change.

Switching to Oxlo.ai looks like this in Python:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key="your-oxlo.ai-api-key"
)

response = client.chat.completions.create(
    model="Llama-3.3-70B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain the trade-offs between tensor parallelism and pipeline parallelism."}
    ],
    stream=True
)

for chunk in response:
    print(chunk.choices[0].delta.content or "", end="")

Oxlo.ai provides streaming responses, function calling, JSON mode, and vision input without requiring you to configure a single GPU driver. There are no cold starts on popular models, so latency remains consistent from the first request of the day.

Hybrid Deployment Patterns

Neither extreme is correct for every organization. A hybrid architecture often delivers the best balance of control and operational simplicity. Teams frequently self-host small, latency-sensitive models on edge or private cloud instances while offloading large reasoning tasks, long-context summarization, and image generation to managed APIs.

For example, a retrieval-augmented generation pipeline might use a locally embedded model for vector search, then call Oxlo.ai for the final generation step with a 128K context window. Because Oxlo.ai pricing is request-based, you can send large retrieved contexts without watching token counters. This is particularly effective for agentic workflows that chain multiple tool calls and reasoning steps, where input length grows quickly.

Decision Framework

When choosing your deployment strategy, evaluate these factors honestly:

  • Traffic pattern: Steady, high-QPS loads often amortize the fixed cost of reserved GPU instances. Spiky or exploratory workloads favor managed APIs.
  • Context length: Long inputs and multi-turn conversations inflate token-based bills. A request-based provider such as Oxlo.ai removes that variable.
  • Compliance: Strict data residency requirements may force on-premise or private cloud deployment. For everything else, a managed API accelerates time to market.
  • Team expertise: Maintaining inference clusters requires platform engineering capacity. If your team is lean, offload the infrastructure.

Conclusion

Cloud platforms provide the substrate for modern AI workloads, but the path from a GPU instance to a reliable production endpoint is longer than most roadmaps assume. Self-hosting delivers maximum control at the cost of operational complexity, while managed APIs like Oxlo.ai remove the infrastructure layer entirely. For teams building with long-context models, agentic systems, or variable traffic, Oxlo.ai request-based pricing and OpenAI-compatible endpoints provide a production-ready foundation without the cluster management burden.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.