Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
I run local models on one home machine: an RTX 5060 with 8 GB, a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. When NVIDIA published Qwen3.8-Flash-Next in NVFP4 (133 GB on disk), the o
I run local models on one home machine: an RTX 5060 with 8 GB, a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. When NVIDIA published Qwen3.8-Flash-Next in NVFP4 (133 GB on disk), the obvious answer was "it doesn't fit". It does now: it runs as a normal chat model in Open WebUI at 9β12 tokens/s, and its output matches the Hugging Face reference implementation.
Code: https://github.com/helgard-orlm/qwen-flash-next-8gb
Who did what, up front: I set the goal, chose the model and made the calls along the way. The engine, kernels and server were written by Claude (Anthropic) in my sessions; the CUDA-graph speed-up was written by Codex (OpenAI) in a parallel session, and Claude then found and fixed a crash in it. "We" below means the three of us.
This post is about the method, the measurements, and what didn't work.
Why this model, of all models
Most of the weights of a mixture-of-experts model are experts, and each token uses only a few of them. So the question is not "how big is the model" but how many expert bytes one token needs.
Qwen3.8-Flash-Next: 48 layers, 512 experts per layer, top-10 routing, and each expert is tiny β three 2560Γ640 matrices, 2.70 MiB in NVFP4 with scales.
10 experts Γ 48 layers Γ 2.70 MiB β 1.27 GiB per token
For comparison, DeepSeek-V4.1-Flash, which we run on the same box with the same trick, needs about 4.2 GiB per token. Same disk, a third of the bytes.
Everything else (embeddings, attention, the 36 Gated DeltaNet layers, router, shared expert, lm_head) is about 7.2 GB in BF16. Storing the big DeltaNet matrices in FP8 with one scale per row brings the resident part to 6.13 GiB of VRAM. The 63 GiB of experts live on the NVMe drive.
The engine
There was no runtime for this architecture that could stream experts, so it's a from-scratch PyTorch + Triton engine. The math was ported from transformers 5.18 (qwen4_exp): DeltaNet, full attention with a block indexer that kicks in beyond ~2048 tokens, four hyper-connection streams, and an n-gram embedding table (PLE, used in layer 1) with 51 billion parameters.
The pieces, in the order they mattered:
1. Repack experts for the disk. In the original files an expert's six tensors are scattered. We rewrote them into one 2,764,800-byte block per expert (675 pages of 4 KiB): gate | up | down | gate_scale | up_scale | down_scale. One expert = one pread with O_DIRECT into page-aligned pinned memory = one copy to the GPU.
2. A RAM cache in front of the disk. LRU over pinned memory (16 GB holds about a quarter of all experts), a pool of reader threads, and two banks of GPU slots so the next layer can load while this one computes.
3. Prefetch into the gap. After a layer's required reads are issued, guess the next layer's experts and read two of them while the GPU is busy (91% of guesses are used). The first version started the prefetch at the same time as the required reads β they shared the disk and everything got slower. It has to go into the gap, not next to the real work.
4. Own kernels. An NVFP4 SwiGLU expert kernel for a single token using Blackwell's hardware cvt.rn.f16x2.e2m1x2 instruction (115 Β΅s per layer, 3.6e-7 from a torch reference), and an FP8 row-scaled matrix-vector kernel. The generic unpack path for the FP8 DeltaNet matrices had been costing ~106 ms per token on its own.
5. Don't read 66 GB per prompt chunk. The first prompt path processed 512 tokens at a time, and each chunk needed nearly every expert, so the whole expert set was read again for every chunk. Running the whole prompt through each layer at once, with experts read in batches of 32, took a 2451-token prompt from 71.6 s to 23.3 s.
6. Read only the PLE rows you need. The n-gram table is a 53.7 GB file; each token needs a handful of rows. Random reads from NVMe: 0.3β0.5 ms per token. (On the HDD it would be a disk seek per row β a second AI acting as reviewer caught that the table was still on the HDD in the first run.)
Where the time actually goes
With every expert already in RAM, a token took 80β88 ms. Splitting it:
- 39 ms β Python issuing GPU work (the GPU was waiting on Python, not the other way round);
- 47 ms β moving 1.33 GB of experts over PCIe (measured 28.6 GB/s, PCIe 5.0 Γ8);
- and the two ran one after the other.
There are 52 CPUβGPU syncs per token (mostly reading the router's chosen experts). We expected them to be the problem; measured, they are cheap. They only hurt because the expert copy for layer L can't start until layer L's router has answered. That pointed to three changes:
- Speculative copies. Run layer L+1's router on layer L's MoE input. It picks the right top-10 65% of the time β enough to start most copies one layer early. (Prefetching top-16 to raise the hit rate moved more bytes over the bus than it saved.)
- CUDA graphs for the 36 one-token DeltaNet layers (Codex). One shared memory pool: 36 private pools ran out of memory on an 8 GB card. Graphs are dropped before any prompt β₯1024 tokens.
- Pin the decode thread to a P-core. On a hybrid CPU the unpinned thread wandered between P and E cores: 93β108 ms per token, jumping around. Pinned to core 0: a flat 81.6 ms.
Results
| configuration | tokens/s |
|---|---|
| first version, experts from disk, no cache | 2.68 |
| disk only + prefetch | 4.72β4.75 |
| RAM cache 16 GB + prefetch | 8.1 |
| live chat, 1000-token answer, before the last three changes | 9.71 |
| live chat, 1000-token answer, after | 11.65 |
Prompt processing: 24 tokens in 4.5 s, 2451 in 23.3 s, 7182 in 98 s. Follow-up messages are fast because the server snapshots the whole recurrent state (DeltaNet matrices + KV + indexer keys) in RAM and only processes the new tail.
How we know it's right
Speed without correctness is a random number generator, so every change was checked against the transformers reference fed with the same quantized weights:
- next-token argmax: 32/32, mean |Ξlogit| 0.077;
- a 2451-token prompt, long enough for the attention indexer to start selecting blocks: 11/11, and the needle-in-a-haystack fact comes back right;
- perplexity EN 2.21 / RU 2.22 β plus control runs that must break: with nibbles swapped it's 2016 / 33,269, with experts removed 588 / 5,351. That checks the NVFP4 unpacking independently of the reference.
What didn't work
- A fused DeltaNet kernel (one kernel for the recurrence, 4.3 β 0.43 ms per token across 36 layers). 5 of 256 tokens came out different from the reference. Rejected β 4 ms isn't worth a model that says different things.
-
The 8192-token context didn't actually work, in the old version either: a 7182-token prompt ran out of VRAM in attention. Fixed by sizing prompt pieces so that
piece Γ (position + piece) β€ 3072Β². -
A crash only after long prompts: CUDA graph capture in the default
globalmode was invalidated by expert-reader threads callingevent.synchronize().capture_error_mode="thread_local"fixed it.
Try it
The repository has the engine, the server (OpenAI-compatible, streaming, a thinking-mode model id), the repack and reference-check tools, and a setup.sh that prepares a working directory from a Hugging Face snapshot. Before publishing, we ran it the way a stranger would: a fresh git clone of the repo on the same box, setup.sh into an empty directory (15 min 48 s, mostly reading ~120 GB from an HDD; spot check of 204 repacked experts against the originals: 0 differences), then the server from the clone. The repacked experts and the extracted non-expert weights came out byte-identical to the production directory, and three greedy test prompts gave identical answers from both.
You need a Blackwell GPU (the NVFP4 kernel uses the sm_120a e2m1 instruction), a fast NVMe drive with ~130 GB free, and patience for a first-time setup that reads most of the 133 GB.
The bigger point: on a mixture-of-experts model, "does it fit in VRAM" is the wrong question. The right one is bytes per token, and a home PC with a fast drive can serve a lot more of them than it looks.
Engine, kernels, server and this write-up: Claude (Anthropic). CUDA graphs and the speculation/pinning A/B: Codex (OpenAI). Goal, model choice and decisions: me. All numbers are from logs on the machine described above.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.