Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 5 min read

Run AI Locally on Your Mac with Ollama: Chat, Tools, RAG and a Custom Model (Real Numbers from a Mac mini M4)

I wanted to know how far a plain Mac mini M4 with 16 GB can go with local AI. No cloud, no API key, nothing leaving the machine. So I installed Ollama and built five small projects with it: a chat app, the OpenAI SDK poi

I wanted to know how far a plain Mac mini M4 with 16 GB can go with local AI. No cloud, no API key, nothing leaving the machine. So I installed Ollama and built five small projects with it: a chat app, the OpenAI SDK pointed at localhost, tool calling, RAG over my own notes, and a custom model.

Everything below is from that real run. The full step-by-step video is here:

What Ollama is (in one paragraph)

Ollama is a small app that downloads open models (Llama, Qwen, Gemma and more) and runs them on your own computer. It serves them on http://localhost:11434, so anything that can make an HTTP request can use them: the terminal, Python, or any tool that speaks the OpenAI API. On Apple Silicon it runs the model on the GPU through Metal.

Install and first model

Download Ollama from ollama.com, drag it to Applications and open it. Then:

ollama pull qwen3:4b-instruct
ollama run qwen3:4b-instruct --verbose

--verbose prints the speed after each answer. On the M4 I got:

Model Speed Notes
Qwen 3 ยท 4B instruct 35.8 tokens/s 100% GPU, 3.2 GB loaded, 4,096-token context
Llama 3.2 ยท 3B 44.8 tokens/s faster, but it didn't know what Ollama is (training cutoff)

Two commands I use all the time:

ollama ps          # what is loaded right now, and on CPU or GPU
ollama show qwen3:4b-instruct   # context length, parameters, template

A loaded model stays in memory for 5 minutes after the last request, then Ollama unloads it.

1. A streaming chat app with memory

The official ollama Python package is the shortest path:

import ollama

messages = [{"role": "system", "content": "You are a friendly tutor. Keep answers short."}]

while True:
    question = input("you โ€บ ")
    if question in {"exit", "quit"}:
        break
    messages.append({"role": "user", "content": question})

    reply = ""
    for chunk in ollama.chat(model="qwen3:4b-instruct", messages=messages, stream=True):
        print(chunk.message.content, end="", flush=True)
        reply += chunk.message.content
    print()
    messages.append({"role": "assistant", "content": reply})  # memory: the model sees the whole chat

"Memory" here is nothing magic: you send the whole conversation back every time.

2. Use the OpenAI SDK, but local

Ollama exposes an OpenAI-compatible endpoint, so existing code works by changing one line:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")  # the key is ignored
r = client.chat.completions.create(
    model="qwen3:4b-instruct",
    messages=[{"role": "user", "content": "Explain an API in one sentence."}],
)
print(r.choices[0].message.content)

This is the easiest way to move a small project off a paid API for testing.

3. Tool calling: the model runs real Python functions

You pass plain Python functions with docstrings. The model decides which ones to call:

import ollama, subprocess

def get_disk_free() -> str:
    """Get the free space on this Mac's main disk."""
    out = subprocess.run(["df", "-g", "/"], capture_output=True, text=True).stdout.split("\n")[1].split()
    return f"{out[3]} GB free of {out[1]} GB"

messages = [{"role": "user", "content": "How much free disk space does my Mac have?"}]
response = ollama.chat(model="qwen3:4b-instruct", messages=messages, tools=[get_disk_free])

for call in response.message.tool_calls or []:
    result = get_disk_free()
    messages += [response.message, {"role": "tool", "content": result, "tool_name": call.function.name}]

print(ollama.chat(model="qwen3:4b-instruct", messages=messages).message.content)

In my run, with a second function for memory, the model called both and answered: 150 GB disk free, 61% of 16 GB memory free.

4. RAG: answers from your own documents

This was the most useful part. I asked plain Llama 3.2 which Mac my studio uses. It had no idea, of course. Then I gave it my notes with RAG:

import numpy as np, ollama
from pathlib import Path

chunks = [p.strip() for f in Path("docs").glob("*.md") for p in f.read_text().split("\n\n") if p.strip()]
vectors = np.array(ollama.embed(model="nomic-embed-text", input=chunks).embeddings)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)

question = "Which Mac does the studio use?"
q = np.array(ollama.embed(model="nomic-embed-text", input=question).embeddings[0])
best = np.argsort(vectors @ (q / np.linalg.norm(q)))[::-1][:3]

context = "\n\n".join(chunks[i] for i in best)
print(ollama.generate(model="llama3.2:3b",
      prompt=f"Answer using only this context:\n\n{context}\n\nQuestion: {question}").response)

The top match scored 0.869, and the same small model answered correctly: "Apple M4 with 16 GB". Same model, same Mac, just the right facts in the prompt.

5. Your own model with a Modelfile

A Modelfile doesn't train anything. It packages a base model with your settings and system prompt:

FROM llama3.2:3b

PARAMETER temperature 0.4
PARAMETER num_ctx 8192

SYSTEM """
You are Rainy Tutor, a patient coding teacher for beginners.
Explain with one simple everyday analogy.
Answer in under 60 words.
"""
ollama create rainy-tutor -f Modelfile
ollama run rainy-tutor "What is a variable?"

Bonus: fine-tuning with MLX, then running it in Ollama

The video also covers a real LoRA fine-tune of Llama 3.2 1B with Apple's MLX: 120 examples, 300 steps in about 6 minutes, 0.456% of the parameters trained. Validation loss was best at step 100 (1.382) and got worse after that (1.811 at step 300), a clear case of overfitting, so I used the step-100 checkpoint. After mlx_lm.fuse, ollama create imported it straight from safetensors at 36.7 tokens/s. One gotcha: it needed the Llama 3 chat TEMPLATE in the Modelfile, otherwise the answers fell apart.

Is a 16 GB Mac enough?

For 3Bโ€“4B models, yes, comfortably: 35โ€“45 tokens/s is faster than you can read. Small models are great for chat, tools, RAG over your own files and private experiments. They are not a replacement for the biggest cloud models on hard reasoning, and a 1B fine-tune will still get facts wrong. Know which job you're giving it.

If you want to see every step on screen (install, each script running, the fine-tune), the full video is on my YouTube channel, and the advanced follow-up fixes a broken RAG search and imports a GGUF model.

What are you running locally right now? I'm picking the next model to test from the comments.

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.