Qwen3.6 27B: how much VRAM does it really need
A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.
In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.
en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original
The Qwen3.6 27B was released in April 2026 as a dense model, Apache 2.0 license, with native 262 thousand tokens of context and multimodal input. Dense 27B usually means “forgets, doesn’t fit on your board.”
But when opening its config.json, the numbers don’t add up as expected. It uses half of the KV cache of a Llama 3 8B — a model three and a half times smaller.
The reason lies in a line of configuration that almost no one reads.
Qwen3.6-27B · from the official config.json
- Parameters
- 27.8 B (dense)
- Layers
- 64
- Full attention
- only 16 of them
- KV heads
- 4 (GQA 24:4)
- Head dimension
- 256
- Native context
- 262,144 tokens
The line that changes everything
In the config.json, among the common keys, there is this:
"full_attention_interval": 4,
"layer_types": ["linear_attention", "linear_attention", "linear_attention", "full_attention", ...]
The model has 64 layers, but only one out of every four uses full attention. The other 48 use linear attention.
This distinction is the most important thing in this post because they behave in opposite ways in memory:
- Full attention stores a key-value pair per token, in each layer. It’s the KV cache, and it grows indefinitely as the conversation progresses.
- Linear attention maintains a state of fixed size, compressing the past into a summary of constant dimension. Double the context and this state does not change in size.
In other words: 48 out of the 64 layers simply do not participate in memory growth. Whoever applies the usual formula — the one from the previous post about VRAM — over the 64 layers will be wrong by a factor of four.
The real KV cache
The formula remains the same, changing only how many layers are included in it:
bytes por token = 2 × camadas_de_atenção_plena × cabeças_kv × dim_cabeça × bytes
2 × 16 × 4 × 256 × 2 = 65.536 bytes = 64 KiB por token
Sixty-four KiB. For comparison, with the same math:
| Model | Layers with KV | KiB/token | 128k context (GB) |
|---|---|---|---|
| Llama 3 8B | 32 out of 32 | 128 | 17.18 |
| Qwen3.6 27B (real) | 16 out of 64 | 64 | 8.59 |
| Qwen3.6 27B if it were dense | 64 out of 64 | 256 | 34.36 |
A model with 27 billion parameters and half the context cost of one with 8 billion. It’s counterintuitive, and it’s the reason why this model fits where it shouldn’t fit.
The weights
With 27.8 billion parameters:
| Quantization | Weights (GB) |
|---|---|
| BF16 | 55.6 |
| Q8_0 | 29.5 |
| Q6_K | 22.9 |
| Q5_K_M | 19.8 |
| Q4_K_M | 17.0 |
Does it fit on your board?
Here’s where the math meets hardware. Weights in Q4_K_M, plus the KV cache of the context, plus 1 GB of overhead for runtime:
| Machine | Memory (GB) | Remaining for context (GB) | Possible context |
|---|---|---|---|
| RTX 5070 · 12 GiB | 12.9 | — | does not fit |
| RTX 5070 Ti · 16 GiB | 17.2 | — | missing 0.8 GB |
| RTX 4090 · 24 GiB | 25.8 | 7.7 | 118k |
| RTX 5090 · 32 GiB | 34.4 | 16.3 | 249k |
| MacBook Pro M4 Max · 48 GiB | 51.5 | 33.5 | 262k (ceiling) |
| Mini PC · 96 GiB | 103.1 | 85.1 | 262k (ceiling) |
Look at the second line. The 5070 Ti is out by 0.8 GB — less than 5% of the board’s memory. The weights alone occupy 17.03 of the 17.18 GB it has; there are 152 MB left for runtime and context, which isn’t even enough to load.
It’s the kind of margin that makes someone buy the wrong board, and the trap has two parts. The unit deceives in favor of the board: the “16 GB” on the box is 16 GiB, or 17.18 decimal GB — and suddenly a model “of 17 GB” seems to fit. The weights do fit. What doesn’t fit is what comes after them: the KV cache and runtime overhead are not optional, and they don’t appear on any label.
What you can reproduce now
The arithmetic part of this post you can check yourself without downloading 17 GB:
# The official configuration — the source of everything above
curl -sL https://huggingface.co/Qwen/Qwen3.6-27B/raw/main/config.json \
| python3 -c "
import json,sys
from collections import Counter
t = json.load(sys.stdin)['text_config']
tipos = Counter(t['layer_types'])
plenas = tipos['full_attention']
kv = 2 * plenas * t['num_key_value_heads'] * t['head_dim'] * 2
print('camadas:', t['num_hidden_layers'], dict(tipos))
print('KV por token:', kv, 'bytes =', kv // 1024, 'KiB')
"
And to check on your own board after downloading:
ollama pull qwen3.6:27b
# With declared context — the runtime pattern is not what you will use
OLLAMA_CONTEXT_LENGTH=131072 ollama run qwen3.6:27b "ok"
# How much was actually reserved and if there was any left on the CPU
ollama ps
nvidia-smi --query-gpu=memory.used,memory.total --format=csv
What I still don’t know
Being explicit about the boundary between what this post proves and what it only suggests.
Proven here: the hybrid architecture, the KV cache of 64 KiB/token, and what the arithmetic says about each machine. All derived from the official config.json, reproducible by the command above.
Not yet measured: the state of linear layers, real GPU consumption after loading, tokens per second on each machine, and what linear attention costs in quality in a 200 thousand token context — because compressing the past into a fixed state is not free, and it’s exactly the test that matters.
That last one is the good question. A model that promises 262k of context and fits on a 4090 is news; if it remembers what was in token 30 thousand when it reaches 200 thousand is another story, and no one answers that with arithmetic.
In the next post I measure it, on six machines. The math says what fits; only measurement says what’s worth it.
Sources: official config.json on Hugging Face · model announcement · qwen3.6:27b on Ollama