Skip to content
Fantástico Mundo de Jon
RSS

When the model doesn't fit: the real cost of offloading to CPU

In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by qwen3:32b-fast, reviewed by the author. Read the Portuguese original

In the post about how much VRAM an LLM occupies I ended with a list of ways out for when the model doesn’t fit. The fourth was moving layers to the CPU, with this caveat: “it works, almost anything fits — and it’s ten times slower.”

Ten times was an estimate. I went to measure.

I was wrong — on the lower side. The same model, on the same machine, with the same question: 66.3 tokens per second on the GPU where it fits, 3.1 on the GPU where it doesn’t. Twenty-one times.

The bench

Bench · August 2026

GPUs
RTX 5090 (32 GiB) and RTX 5070 (12 GiB), same machine
Runtime
Ollama 0.32.5, flash attention enabled
KV cache
q8_0 — the recommended swap in the previous post
Context
32.768 tokens
Load
real technical prompt, 200 tokens of output
Sampling
median of 3 runs, after warmup

Six models, from 4.9 to 20 GB, chosen to cross the 12 GiB boundary of the smaller GPU. Each runs in its own runtime instance, with the GPU fixed — for reasons that become clear at the end of the post.

The measurement

RTX 5090 · 32 GiB

Model File On GPU tok/s TTFT Prefill tok/s
llama3.1:8b 4.9 GB 100% 215.3 0.20 s 6946
qwen3-vl:8b 6.1 GB 100% 190.1 0.20 s 6320
gemma3:12b 8.1 GB 100% 115.1 0.44 s 2828
deepseek-r1:14b 9.0 GB 100% 125.6 0.21 s 3889
gpt-oss:20b 13.0 GB 100% 230.1 0.36 s 6036
qwen3:32b 20.0 GB 100% 66.3 0.26 s 2164
Median of 3 runs on RTX 5090, Ollama 0.32.5, KV cache q8_0, 32k context, 200 tokens of output. TTFT measured on the client; throughput reported by the runtime. All six models fit entirely in VRAM.

On the 32 GiB GPU everything fits, and the reading is simple: a larger model tends to be slower. Now the same battery on the 12 GiB GPU.

RTX 5070 · 12 GiB

Model File On GPU tok/s TTFT Prefill tok/s
llama3.1:8b 4.9 GB 100% 105.9 0.26 s 3864
qwen3-vl:8b 6.1 GB 100% 100.7 0.21 s 3766
gemma3:12b 8.1 GB 100% 62.5 0.48 s 1562
deepseek-r1:14b 9.0 GB 100% 59.6 0.26 s 2141
gpt-oss:20b 13.0 GB 75% 52.2 0.84 s 407
qwen3:32b 20.0 GB 52% 3.1 3.62 s 46
Same battery, same machine, RTX 5070. The 'On GPU' column comes from the runtime itself and indicates how much of the model remained resident in VRAM; the rest was moved to system RAM. Confirmed against the VRAM peak read from the driver.

The first four models fit entirely and the larger GPU delivers about double — a real difference, but linear and dull. The story is in the last two lines.

What the drop has of specificity

The qwen3:32b doesn’t become “twice as slow” on the smaller GPU. It becomes 21 times slower, and the degradation isn’t uniform:

Metric RTX 5090 (100% GPU) RTX 5070 (52% GPU) Factor
Generation 66.3 tok/s 3.1 tok/s 21×
Time to first token 0.26 s 3.62 s 14×
Prompt processing 2164 tok/s 46 tok/s 47×
qwen3:32b on both GPUs, same question and same runtime. The prefill is the most affected stage: it's where the model reads what you wrote before starting to respond.

The prefill is the most impacted, with a factor of 47. It’s the stage where the model processes what you wrote — and it’s precisely the stage that grows when the prompt is long. In other words: the more context you use, the more painful the offload becomes, exactly in the scenario where having a large model would make a difference.

It’s worth noting that the 3.1 tokens per second came from an already optimized configuration: KV cache in q8_0, which is the second recommendation from that list. Without it, the number would be worse.

None of this causes an error

This is the point that made me write the post instead of just updating a table.

When the model doesn’t fit, nothing breaks. There’s no error, no warning, no red line in the log. The runtime moves part of the layers to RAM and keeps working. You ask, it answers, the answer is correct.

It just takes twenty-one times longer.

And since you probably never ran this model on a GPU where it would fit, you have nothing to compare to. The natural conclusion is “a large model is slow on this machine” — when the correct one would be “this specific model is half in RAM, and there’s a better model for this GPU”.

The previous recommendation is confirmed

The previous post’s list ended like this: “a well-chosen 12B model running entirely on the GPU delivers more value per second than a 32B model half on the CPU.”

In the numbers from this bench, on the same 12 GiB GPU:

  • gemma3:12b, entirely on the GPU: 62.5 tok/s
  • qwen3:32b, 52% on the GPU: 3.1 tok/s

Twenty times faster. The recommendation was correct, and by a margin larger than I imagined when I wrote it.

Two things the arithmetic didn’t predict

The file size doesn’t say if it will fit. The gpt-oss:20b occupies 13 GB on disk — more than the 12 GiB nominal capacity of the GPU — and yet it loaded 75% on the GPU and maintained 52.2 tok/s, perfectly usable. Meanwhile, the qwen3:32b, with 20 GB, ended up at 52%. Quantization, weight format, and cache configuration change the math, and none of that appears in the size shown by ollama list.

A larger model isn’t a slower model. The fastest on the RTX 5090 wasn’t the smallest in the list: it was the gpt-oss:20b, with 230.1 tok/s — ahead of a model of 8B with less than half the size. It’s an MoE architecture, which activates only a fraction of the parameters per token. Choosing a small model “because it’s faster” is using a rule that doesn’t hold up to measurement.

The mistake that almost made it into this post

Here’s what I would do differently — and the reason each model runs in an isolated instance with the GPU fixed.

The first version of this bench produced these numbers for the smaller GPU:

Model RTX 5090 RTX 5070 (false) RTX 5070 (real)
qwen3:32b 65.8 tok/s 61.6 tok/s 3.1 tok/s
llama3.1:8b 215.7 tok/s 186.8 tok/s 105.9 tok/s
Middle column: result from the battery where the GPU wasn't actually fixed. The numbers are plausible and completely wrong — they were measured on the larger GPU.

Notice the middle column. 61.6 against 65.8 tokens per second seems exactly the difference you’d expect between two different GPUs. Nothing there raises suspicion.

Except the RTX 5070 was never touched. The entire battery ran on the 5090.

The reason: on a machine with two GPUs, CUDA_VISIBLE_DEVICES isn’t enough to fix the GPU in Ollama. It filters the CUDA backend — but Ollama also enumerates GPUs through the Vulkan backend, which is a different API and ignores that variable. The scheduler saw the “hidden” GPU, noticed it had more free memory, and chose it. The log makes this explicit for those who know what to look for:

grep 'selecting single GPU' ollama.log

msg=“user overrode visible devices” CUDA_VISIBLE_DEVICES=1 library=Vulkan name=Vulkan0 description=“NVIDIA GeForce RTX 5090” library=CUDA name=CUDA0 description=“NVIDIA GeForce RTX 5070” msg=“selecting single GPU” main_gpu=0 library=Vulkan description=“NVIDIA GeForce RTX 5090”

saída

What exposed the error wasn’t the number — it was instrumenting the VRAM peak read from the driver. The GPU I thought I was measuring showed 15 MB in use. An idle GPU doesn’t generate 61 tokens per second.

The fix are three variables, not one:

CUDA_VISIBLE_DEVICES=1      # filtra o backend CUDA
GGML_VK_VISIBLE_DEVICES=1   # filtra o backend Vulkan
OLLAMA_VULKAN=0             # ou simplesmente desliga o Vulkan

The other four, for those who want to measure their own machine:

  1. The first generation measures your SSD. It includes loading the model from disk. Without discarding it with a warmup, the “response time” of a 20 GB model becomes a storage test.
  2. Unloading the model between repetitions ruins the TTFT. With keep_alive=0, each repetition reloads everything, and the time to the first token becomes reload time: I measured 8.5 s where the real value was 0.19 s.
  3. Repeating the same prompt makes the cache lie. The runtime reuses the already processed prefill and the TTFT drops. Vary the text with each repetition.
  4. Reasoning models don’t respond through the field you’re reading. qwen3 and deepseek-r1 emit the thought in a separate field from the response. Timing the first token by looking only at response registers TTFT zero in these models — and zero isn’t a suspicious value, it’s an impossible value.

Measure yours

The bench code is open, and runs on any machine with an NVIDIA GPU and Ollama. It runs its own runtime instance, one GPU at a time, without touching the machine’s configuration:

git clone https://github.com/SEU-USUARIO/bench-llm-local
cd bench-llm-local && pip install requests matplotlib
python3 bench.py       # todas as GPUs detectadas
python3 relatorio.py   # tabela e gráfico

A warning to close, because it’s more important than any number above: speed isn’t competence. The gpt-oss:20b was the fastest in this bench, and that says nothing about it being more accurate than the qwen3:32b, which runs at a third of the speed. If the model gets the task wrong, tokens per second just mean reaching the wrong answer faster.

What this measurement answers is the other half of the question — the one that’s often skipped, and the one that makes people give up on running a local model thinking the machine can’t handle it. Almost always it can. Just not with that model.