Dev.to AI 🤖 Ai 👁 0 📖 4 min read

GPU offload says Max, but only 54 of 65 layers loaded: how to check where your model actually runs

A 27B model at Q4 was generating 12-18 tokens per second on an RTX 3090. Same model, same settings a week earlier, it was doing 60-70. Nothing in the app looked broken. The model config still said GPU Offload: Max. That

A 27B model at Q4 was generating 12-18 tokens per second on an RTX 3090. Same model, same settings a week earlier, it was doing 60-70. Nothing in the app looked broken. The model config still said GPU Offload: Max.

That's the trap. Max doesn't mean all layers on the GPU. It means as many as the app's size estimate thinks will fit. When that estimate comes out pessimistic, the loader quietly leaves layers on the CPU, and you pay for it on every token instead of noticing it once.

Here's how to check what actually happened, and what to try.

The load log is the source of truth, not the dropdown

LM Studio writes a log per model load. They live under your user profile in .lmstudio. The paths I've seen quoted are %USERPROFILE%\.lmstudio\logs\ for the app log and %USERPROFILE%\.lmstudio\server-logs\ for the per-load server log, so have a poke around if yours sit somewhere else.

Load a model, open the newest log, and look for this pair:

Limit weight offload to dedicated GPU Memory: ON
Model load size estimate with adjusted num offload layers '54'

"Adjusted" is the word doing the damage. The setting asked for Max. The estimator adjusted it down to 54 because its model of your VRAM said the rest wouldn't fit.

Then llama.cpp prints the final answer on its own line:

offloaded 54/65 layers to GPU

If both numbers match, everything landed. If they don't, you're running a hybrid and the CPU is handling part of every token. On a 27B that's the difference between roughly 45 tokens per second and 15.

The check that doesn't need logs

Watch GPU and CPU utilization while the model generates.

Full offload looks like GPU utilization pinned high (the reporter measured around 97%) with the CPU mostly idle. Partial offload looks like low, spiky GPU utilization, a busy CPU, and throughput down by 2-3x. If it looks like the second case while the settings say Max, the estimator is your suspect.

Why the estimate trims layers at all

The machinery shows up clearly in another report in the same tracker. With the VRAM cap armed, the estimator runs and adjusts. With the cap off, it skips the adjustment step entirely:

[LM Studio] Model load size estimate with raw num offload layers 'max' and context length '8192'
[LM Studio] Strict GPU VRAM cap is OFF: GPU offload layers will not be checked for adjustment
[LM Studio] Resolved GPU config options:   Num Offload Layers: max

So the setting behind that first log line, something like "Limit weight offload to dedicated GPU memory" in the load settings, is what arms the trimming. It isn't a silly feature. It stops you from requesting more layers than the card can hold and getting worse performance for it. It just isn't always right.

What to try, in order

  1. Switch the runtime backend. This was the change that reliably worked for the reporter. The engine was set to the CUDA 12 runtime; switching to the plain CUDA runtime, or to Vulkan, restored full offload. Throughput went from 12-18 tok/s to 41.9-45.4 tok/s at roughly 97% GPU utilization, across four different 27B models and a range of CUDA 12 runtime versions. A dropdown change for about a 3x speedup.
  2. Set an explicit layer count instead of Max, so nothing is left to the estimator. Unload and reload, then read the log line again. If you set 65 and the log still says 54, the trimming is happening anyway and you're back to step 1.
  3. Reduce the context length. The KV cache for a long context eats into the budget the estimator is working with, so a shorter context can leave more room for weights. Worth trying, though the reporter found it didn't reliably stop the behaviour.
  4. From the CLI: lms load <model> --gpu max (or off, or a fraction like 0.5), --context-length to cap context, and --estimate-only to see what the loader plans before you commit to the load. Those are in LM Studio's CLI docs.

The honest part

The CUDA 12 trimming is an open issue with no official fix as of late September 2026, so everything above is diagnosis plus a workaround, not a patch. Nobody has confirmed why the CUDA 12 path over-trims while the regular CUDA and Vulkan runtimes don't. The running theory in the thread is that the CUDA 12 build reserves more memory for graph and KV buffers, which makes the estimate come out larger, but that's a commenter's theory rather than a maintainer's finding. Treat it as a hypothesis.

Worth saying out loud too: partial offload is sometimes the correct outcome. If your card genuinely can't hold the model, turning the cap off doesn't create VRAM. It just lets you overcommit, and overcommitting usually performs worse than a clean split.

One related pattern, because it's the same class of problem on AMD hardware: an RX 6900 XT (RDNA2) on Windows can fail outright with "RDNA2 not supported with the ROCm version you are using" when the ROCm engine gets selected, while Vulkan works fine on the same card. Different message, same lesson. The backend picker matters more than the offload slider, so bisect the runtime before changing anything else.

If you remember one thing from this: after loading a model, grep the load log for the offload count. Max is a request, not a receipt.

The source report, with the full logs, is here: https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/2312

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.