Dev.to AI 🤖 Ai 👁 0

Complete Guide to Running AI Models Locally for Free

Run powerful AI models on your own hardware for free. No API keys needed. Install Ollama curl -fsSL https://ollama.ai/install.sh | sh ollama pull qwen2.5:0.5b ollama pull gemma:2b Use via AP

Run powerful AI models on your own hardware for free. No API keys needed.

Install Ollama

curl -fsSL https://ollama.ai/install.sh | sh
ollama pull qwen2.5:0.5b
ollama pull gemma:2b

Use via API

import requests

def ask_ollama(prompt, model="qwen2.5:0.5b"):
    resp = requests.post("http://localhost:11434/api/generate", json={
        "model": model,
        "prompt": prompt,
        "stream": False,
    })
    return resp.json()["response"]

Available Free Models

Model Size Best For
qwen2.5:0.5b 400MB Fast responses, simple tasks
gemma:2b 1.5GB General purpose
tinyllama 600MB Lightweight tasks
phi3:mini 2GB Reasoning tasks

Cost Comparison

  • OpenAI API: $0.002/1K tokens
  • Ollama local: $0 (just electricity)
  • For 1M tokens/month: Save $2

Production Tips

  • Use qwen2.5:0.5b for speed
  • Use gemma:2b for quality
  • Set OLLAMA_KEEP_ALIVE=24h to keep models loaded
  • Limit to 1 thread for OS stability

All these models used in the https://create-openings-unsigned-garden.trycloudflare.com storefront.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.