Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 7 min read

An offline voice agent for my terminal in 52 MB (Whistle + Needle)

Two of the biggest Hacker News threads this week were about Whistle, a 16.9 MB speech-to-text model, and DeepSeek V4.1 Flash. I wanted to see what you can build with the first one, so I put a voice front end on my termin

An offline voice agent for my terminal in 52 MB (Whistle + Needle)

Two of the biggest Hacker News threads this week were about Whistle, a 16.9 MB speech-to-text model, and DeepSeek V4.1 Flash. I wanted to see what you can build with the first one, so I put a voice front end on my terminal, with the second as a fallback.

Say "show me the last fifteen minutes of logs for the checkout service" and it prints kubectl logs deploy/checkout --since=15m. Speech recognition and tool choice both run on the CPU. Your audio never leaves the machine. If the local model returns no usable call, the transcript goes to a cloud model for one try.

The stack

  • Whistle (Cactus Compute, Apache-2.0): speech to text in a single 16.9 MB file. It handles English, German, French, Spanish, Italian, Dutch and Polish.
  • Needle, from the same package: a small tool-calling model (121M parameters, per its README). The needle3 file it downloaded for me was 35.3 MB.
  • DeepSeek V4.1 Flash as the fallback, used only when Needle returns no usable call.

The two local model files add up to 52 MB.

Both local models come from one package, pip install cactus-needle, which downloads the weights on first use. Add the [mic] extra for microphone input, and pip install openai for the fallback.

Voice agent pipeline: microphone, Whistle speech to text, Needle tool choice, a confirmation gate, and a DeepSeek fallback

The code

The script has six tools, a confirmation gate and a cloud fallback. Run it with a WAV file as the argument, or with no argument to record four seconds from the mic. It prints each command instead of running it.

import json, os, sys
from typing import Literal
os.environ.setdefault("NEEDLE_TELEMETRY", "0")  # the package sends anonymous usage counts unless this is 0
import needle

Service = Literal["checkout", "billing", "payments", "search", "api"]
Env = Literal["staging", "production"]

@needle.tool
def run_tests(service: Service):
    "Run the test suite of one service."
    return ["pytest", f"services/{service}/tests"]

@needle.tool
def tail_logs(service: Service, minutes: int):
    "Show the recent log lines of a service."
    return ["kubectl", "logs", f"deploy/{service}", f"--since={minutes}m"]

@needle.tool
def scale_service(service: Service, replicas: int):
    "Change how many replicas of a service are running."
    return ["kubectl", "scale", f"deploy/{service}", f"--replicas={replicas}"]

@needle.tool
def deploy(service: Service, environment: Env):
    "Deploy the latest build of a service to an environment."
    return ["./deploy.sh", service, environment]

@needle.tool
def rollback(service: Service, environment: Env):
    "Roll a service back to its previous release in an environment."
    return ["./rollback.sh", service, environment]

@needle.tool
def git_status():
    "Show uncommitted changes in the repository."
    return ["git", "status", "--short"]

TOOLS = {f.__name__: f for f in [run_tests, tail_logs, scale_service, deploy, rollback, git_status]}
RISKY = {"deploy", "rollback", "scale_service"}
WORDS = ["checkout", "billing", "payments", "search", "staging", "production", "replicas"]

ear = needle.Whistle()
brain = needle.Needle(tools=list(TOOLS.values()), stateless=True)

def hear(wav=None):
    if wav is None:
        import sounddevice as sd  # pip install "cactus-needle[mic]"
        wav = sd.rec(4 * 16000, samplerate=16000, channels=1, dtype="float32")
        sd.wait()
        wav = wav[:, 0]
    return ear.transcribe(wav, keywords=WORDS)["text"]

def ask_cloud(text):
    from openai import OpenAI
    client = OpenAI(api_key=os.environ["DEEPSEEK_API_KEY"], base_url="https://api.deepseek.com")
    reply = client.chat.completions.create(
        model="deepseek-flash",
        messages=[{"role": "user", "content": text}],
        tools=[{"type": "function", "function": needle.build_schema(f)} for f in TOOLS.values()],
        extra_body={"thinking": {"type": "disabled"}},
    )
    calls = reply.choices[0].message.tool_calls or []
    return [{"name": c.function.name, "arguments": json.loads(c.function.arguments)} for c in calls], 0.0

def handle(text):
    plan = brain.complete(text)
    calls, sure = plan.get("function_calls") or [], plan.get("confidence") or 0.0
    if plan.get("validation", {}).get("ungrounded"):
        calls = []  # it filled in a value you never said
    if not calls and os.environ.get("DEEPSEEK_API_KEY"):
        calls, sure = ask_cloud(text)
    if not calls:
        print(f"Heard {text!r}. Not sure what to run, say it again.")
    for call in calls:
        name = call["name"]
        if name not in TOOLS:
            continue
        cmd = TOOLS[name](**(call.get("arguments") or {}))
        if (name in RISKY or sure < 0.5) and input(f"{' '.join(cmd)}  (yes to run) ") != "yes":
            continue
        print("would run:", " ".join(cmd))  # swap for subprocess.run(cmd) once you trust it

if __name__ == "__main__":
    text = hear(sys.argv[1] if len(sys.argv) > 1 else None)
    print("heard:", text)
    handle(text)

Four details matter:

  1. Literal types become enums in the tool schema, and Needle's decoder is constrained by the schema, so it can only pick a service that exists. With plain str types it returned Checkout and search service, neither of which is a deployment name.
  2. keywords biases Whistle toward your service names and the other words you expect to say.
  3. The gate. Anything risky, or anything under 0.5 confidence, asks before it runs. Cloud answers carry no confidence score, so they always ask. If validation.ungrounded flags an argument value you never said, the plan is dropped.
  4. Telemetry. By default the package sends anonymous usage counts (function name, versions, OS and a random install ID, no prompts or outputs). The script turns that off with NEEDLE_TELEMETRY=0. Even then, each model load fetches a small config file from Hugging Face, which counts as a download.

What I measured

I ran six spoken commands on a 2 vCPU cloud VM (Intel Xeon at 2.8 GHz, no GPU), five runs each, and took the median. The audio came from Piper TTS, so treat the accuracy part as a smoke test, not a benchmark.

Latency per spoken command on a 2 vCPU VM: speech to text 133 to 290 ms, tool choice 320 to 700 ms

  • Speech to text took 133 to 290 ms for 1.3 to 3.2 seconds of audio.
  • Picking the tool took 320 to 700 ms.
  • Startup, which loads both models, took 2.65 seconds, and the process peaked at 144 MB of memory.

Whistle misheard three of the six commands. "Show me" became "Jomy", "Scale" became "Gale", and "in" became "and". Here is what Needle made of them:

  • "Jomy the last 15 minutes of logs for the Checkout service" became tail_logs(checkout, 15) at 0.81 confidence.
  • "Gale the API service to three replicas" became the right scale_service call plus an extra run_tests, at 0.34 confidence. Both calls stop at the gate, so the extra one does not run unless I say yes.
  • "Run the tests and the payments folder" got no call (0.06 confidence). That is the case the cloud fallback is for.

The three clean transcripts all mapped to the right tool and arguments, at 0.64 to 0.86 confidence.

Cactus reports 11.1 ms to first token for a 10 second clip on an Apple M4 Pro CPU. My numbers cover the whole transcription on a small cloud VM, so the two do not compare directly.

What the fallback costs

The fallback sends the six tool schemas plus your sentence. DeepSeek wraps the schemas in about 210 tokens of its own tool instructions. Counted with DeepSeek's tokenizer in the prompt format published with the model, a short command comes to about 700 input tokens. The tool call that comes back is about 40 tokens with one argument and 56 with two, so call it 50.

At DeepSeek's off-peak rates for deepseek-flash (0.15 USD per million input tokens on a cache miss, 0.60 per million output):

input:  700 x 0.15 / 1,000,000 = 0.000105 USD
output:  50 x 0.60 / 1,000,000 = 0.000030 USD
total per fallback             = 0.000135 USD, about 13.5 cents per 1,000

Peak hours (01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, except Chinese public holidays) cost double, so about 27 cents per 1,000. Thinking mode is on by default for this model, which is why the code sends "thinking": {"type": "disabled"}. A tool call this simple does not need a chain of thought.

Commands handled locally cost nothing per call. For a personal tool, what matters more to me is that the audio stays on the machine and the local path works with no network at all.

Where it breaks

  • Seven languages only.
  • It records a fixed four seconds per run. Whistle has a streaming mode (Whistle.stream) that I have not wired in yet.
  • Needle picks tools. It is not a chat model, so it will not answer questions.
  • I only tested synthetic speech. Real voices, accents and room noise will change the error rate, so test it with your own voice before you trust it near production.

A phone agent is a different problem: it has to listen continuously and answer fast enough to hold a conversation. The production voice agent I built runs on LiveKit at about 2.5 cents a minute. If you are picking speech to text for something real-time, I compared the main options in the best STT for voice agents in 2026, and the AI voice agent page covers what we build for clients.

Sources: Cactus Compute's Whistle announcement and the Cactus-Compute/whistle model card, the cactus-needle README and source, DeepSeek's API pricing and thinking mode docs, and the DeepSeek-V4.1-Flash model card, tokenizer and prompt encoding reference. Latency and accuracy numbers are from my own runs on 10 October 2026.

I used an AI assistant while drafting this.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.