The Architecture I Built vs. The Architecture I Could Deploy
I participated in a competition and it gave me an insight into something interesting and it led me to build Vicquant. An offline financial advsior app but when i was building it, i ran into a problem and and this aid me
I participated in a competition and it gave me an insight into something interesting and it led me to build Vicquant. An offline financial advsior app but when i was building it, i ran into a problem and and this aid me in solving real world problem and be innovative for the time being.
Vicquant's core idea is a dual local-model setup with zero cloud dependency: Llama-3.1-8B-Instruct for deep financial reasoning, Phi-3.5-mini as a fast, low-RAM fallback, both running through llama.cpp directly on the user's device. I will explain the reason for the split in an article soon, the short version is that reasoning quality and response speed pull in opposite directions, and one model can't optimize for both.
That design is real. The code exists, the model-routing logic exists, and it passed testing on an 8GB RAM / 4 vCPU target at roughly 12-15 tokens/second.
What I could actually deploy
Here's the part that doesn't usually make it into these posts: running an 8B-parameter model needs meaningful RAM and CPU, and the PaaS tier I'm currently hosting on have a Pro plan available which i haven't yet purchased. So the version of Vicquant live right now cannot run the local models in production.
Rather than leave the deployment broken or ship nothing, I connected it to OpenRouter so the app can serve AI responses today. Practically, this means: the deployed demo currently requires internet and routes model calls through OpenRouter. Offline mode is off for this deployment. The local-inference architecture isn't gone, it's in the codebase, tested, and it comes back as soon as i can pay for the hosting plans. But I'd rather say that directly than let a live demo imply a capability it isn't currently providing.
Why I'm writing this instead of hiding it
Two reasons.
First, credibility compounds. If someone inspects network requests on the live demo and finds a mismatch with what I claimed, every other technical claim I've made gets discounted too; the backtesting gate, the test suite, all of it. Saying the true thing now costs me one slightly less impressive post. Saying the false thing costs trust I can't easily rebuild.
Second, this constraint is itself worth writing about. The gap between "architecture you design" and "architecture you can afford to deploy on day one" is a real, common problem, and it's one most build-in-public content glosses over in favor of a clean narrative. The clean narrative isn't always the honest one.
What's unaffected by this
The things that don't depend on which model is answering:
147/147 tests passing across indicators (SMA, EMA, RSI, MACD, Bollinger, ATR), the signal engine (candlestick patterns, z-score mean reversion, momentum), and the backtesting engine (2-year minimum history, walk-forward validation, costs and slippage included).
The AI never computes anything. Every number comes from deterministic Python. Whether the explanation comes from a local Phi-3.5 model or an OpenRouter-hosted model, its job is identical: describe an already-correct number in plain language, never generate one.
The backtest gate. No signal reaches a user without proving itself against two years of historical data first, regardless of which model is running.
What's next
Local inference returns once the hosting constraint lifts. I'll post that update when it happens, not before. Thanks for reading
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes â full credit and traffic to the original publisher.