Complete Guide to Running AI Models Locally for Free
Run powerful AI models on your own hardware for free. No API keys needed. Install Ollama curl -fsSL https://ollama.ai/install.sh | sh ollama pull qwen2.5:0.5b ollama pull gemma:2b Use via AP
Run powerful AI models on your own hardware for free. No API keys needed.
Install Ollama
curl -fsSL https://ollama.ai/install.sh | sh
ollama pull qwen2.5:0.5b
ollama pull gemma:2b
Use via API
import requests
def ask_ollama(prompt, model="qwen2.5:0.5b"):
resp = requests.post("http://localhost:11434/api/generate", json={
"model": model,
"prompt": prompt,
"stream": False,
})
return resp.json()["response"]
Available Free Models
| Model | Size | Best For |
|---|---|---|
| qwen2.5:0.5b | 400MB | Fast responses, simple tasks |
| gemma:2b | 1.5GB | General purpose |
| tinyllama | 600MB | Lightweight tasks |
| phi3:mini | 2GB | Reasoning tasks |
Cost Comparison
- OpenAI API: $0.002/1K tokens
- Ollama local: $0 (just electricity)
- For 1M tokens/month: Save $2
Production Tips
- Use qwen2.5:0.5b for speed
- Use gemma:2b for quality
- Set OLLAMA_KEEP_ALIVE=24h to keep models loaded
- Limit to 1 thread for OS stability
All these models used in the https://create-openings-unsigned-garden.trycloudflare.com storefront.
📰 Read the original article on Dev.to AI
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.