The AI That Never Leaves My Laptop
Show almost any enterprise a useful AI workflow and the demo goes well right up until someone from security asks the only question that matters: where does the data go? For every mainstream assistant the answer is some
Show almost any enterprise a useful AI workflow and the demo goes well right up until someone from security asks the only question that matters: where does the data go?
For every mainstream assistant the answer is some version of "to a datacentre we operate, under a policy you should read." For a lot of teams that's fine. For the ones I work with โ regulated industries, client-confidential material, an internal audit function with opinions โ it ends the conversation. Not because the tool is bad, but because the paperwork to say yes costs more than the productivity it buys.
So I built the other thing. OnDevice.ai is a private AI workspace that runs on my own machine: streaming chat, project-based context, file analysis, and real PowerPoint, Excel and Word generation. Inference runs locally through Ollama or LM Studio. Cloud is an option, not an assumption.

Running a 30B model locally on a MacBook. Backend: Ollama ยท Model: Glimmer 30B (GGUF).
Specs before bytes
There's an obvious way to make an AI produce a PowerPoint file: ask the model for one. It's also a bad way. OOXML is a zip of interdependent XML parts with relationship graphs and content-type manifests, and a single malformed relationship gives you a file that refuses to open โ usually when your client double-clicks it.
So the model never touches the bytes:
- The model produces a JSON intermediate spec (
SlideDeckSpec,WorkbookSpec,DocumentSpec), which is schema-validated first. - A fixed, deterministic generator built on python-pptx, openpyxl or python-docx turns the spec into the file.
- Structural validation is a hard gate on download โ zip, OOXML parts, relationship graph, well-formed XML. Fail, and the file is never offered. LibreOffice render checks run too, but only as advisory, in an isolated process on temp copies.
LLMs handle conversation, analysis and planning language. Generators own bytes.
The payoff: a file that downloads is a file that opens. Not "usually."
One boundary, deliberately ordinary stack
Next.js + TypeScript at the front, FastAPI + Python behind it, PostgreSQL for state, Redis optional, local filesystem for binaries. The interesting part is where the boundary sits, and a few principles that explain the refusals as much as the features:
- Narrow harness โ explicit services over plugin ecosystems. Nothing dynamic gets to run.
- Files โ instructions โ uploads are labelled untrusted in the prompt, so a document can't tell the system what to do.
- Keys stay server-side โ no provider secret ever reaches the browser.
Two local runtimes, honestly separated
Ollama runs GGUF; LM Studio on a Mac runs MLX. They aren't interchangeable, and pointing one at the other's model directory gives you confusing failures. So the backends are kept explicitly separate, the model catalogue is filtered by format, and there's a bridge script for the one case where you want a GGUF from your LM Studio library registered into Ollama.
The shipped presets span roughly 12B to 80B parameters across both runtimes, on a MacBook with 64 GB of unified memory โ a good laptop, not a rack.
Search that stays off unless it's needed
A local model will confidently invent this morning's news, so there's optional retrieve-then-generate grounding. In the default auto mode it only searches when the message looks live โ news, prices, "latest", a URL. Ask it to restructure a paragraph and nothing leaves the machine. The honest caveat is written into the architecture doc: when search is on, queries leave the host.
Against Claude and ChatGPT
On raw capability, Claude and ChatGPT win, and it isn't close. More knowledge, better reasoning, far better taste. Anyone claiming a 30B model on a laptop matches them is selling something.
But the commercial services win every row about capability and convenience, and this wins every row about control: where data goes, marginal cost (electricity), offline use, model choice, search behaviour, auditability.
The way I think about it: a frontier assistant has a high ceiling and no floor. This has a lower ceiling and a hard floor. The floor is set by the pipeline and doesn't move. The ceiling is set by whatever model I loaded this morning โ so it rises every time open-weight models improve, without me writing a line of code. Own the floor; let the ceiling arrive by download.
What it deliberately won't do
No computer use, no browser automation, no MCP servers or third-party connectors, no multi-agent orchestration, no user-facing code execution. Every one of those is a capability I'd enjoy having. Every one also widens the blast radius of a very confident text predictor.
You can have a large trust boundary or a small blast radius. Choosing the small one is a position, not a limitation.
Where it stands
It's a live build, not a launch: roughly 19,000 lines, 120 catalogued features, 19 test modules, and a current architecture document. Document generation is the weakest link right now โ a frontier-model deck still looks better, because the content planning comes from a stronger model.
What already works is the part I wasn't sure about at the start: a real AI workspace โ projects, streaming chat, files, artifacts, search when you want it โ on a laptop, with nothing leaving the machine.
Building it was also the best way I've found to learn the whole stack โ auth, streaming, orchestration, generation, validation โ the parts a hosted API keeps hidden. That learn-by-doing approach is what I write about on Curious Bit.
The full build notes, with the architecture diagrams, model table, benchmark caveats and the complete comparison, are here: The AI That Never Leaves My Laptop โ full write-up
Originally published on Curious Bit.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.