Skip to content
Fantástico Mundo de Jon
RSS

Qwen3.6 27B: how much VRAM does it really need

A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original

The Qwen3.6 27B was released in April 2026 as a dense model, Apache 2.0 license, with native 262 thousand tokens of context and multimodal input. Dense 27B usually means “forgets, doesn’t fit on your board.”

But when opening its config.json, the numbers don’t add up as expected. It uses half of the KV cache of a Llama 3 8B — a model three and a half times smaller.

The reason lies in a line of configuration that almost no one reads.

Qwen3.6-27B · from the official config.json

Parameters
27.8 B (dense)
Layers
64
Full attention
only 16 of them
KV heads
4 (GQA 24:4)
Head dimension
256
Native context
262,144 tokens

The line that changes everything

In the config.json, among the common keys, there is this:

"full_attention_interval": 4,
"layer_types": ["linear_attention", "linear_attention", "linear_attention", "full_attention", ...]

The model has 64 layers, but only one out of every four uses full attention. The other 48 use linear attention.

This distinction is the most important thing in this post because they behave in opposite ways in memory:

  • Full attention stores a key-value pair per token, in each layer. It’s the KV cache, and it grows indefinitely as the conversation progresses.
  • Linear attention maintains a state of fixed size, compressing the past into a summary of constant dimension. Double the context and this state does not change in size.

In other words: 48 out of the 64 layers simply do not participate in memory growth. Whoever applies the usual formula — the one from the previous post about VRAM — over the 64 layers will be wrong by a factor of four.

The real KV cache

The formula remains the same, changing only how many layers are included in it:

bytes por token = 2 × camadas_de_atenção_plena × cabeças_kv × dim_cabeça × bytes

2 × 16 × 4 × 256 × 2 = 65.536 bytes = 64 KiB por token

Sixty-four KiB. For comparison, with the same math:

Model Layers with KV KiB/token 128k context (GB)
Llama 3 8B 32 out of 32 128 17.18
Qwen3.6 27B (real) 16 out of 64 64 8.59
Qwen3.6 27B if it were dense 64 out of 64 256 34.36
Arithmetic from each model's config.json with cache in FP16 — not a measurement. The third row is a counterfactual: what the same model would cost if all layers were full attention.

A model with 27 billion parameters and half the context cost of one with 8 billion. It’s counterintuitive, and it’s the reason why this model fits where it shouldn’t fit.

The weights

With 27.8 billion parameters:

Quantization Weights (GB)
BF16 55.6
Q8_0 29.5
Q6_K 22.9
Q5_K_M 19.8
Q4_K_M 17.0
Effective bits per weight of the llama.cpp K-quants. Ollama publishes qwen3.6:27b with 17 GB in Q4_K_M — the calculation matches to the decimal place, which is a good sign that the bits per weight used here are correct.

Does it fit on your board?

Here’s where the math meets hardware. Weights in Q4_K_M, plus the KV cache of the context, plus 1 GB of overhead for runtime:

Machine Memory (GB) Remaining for context (GB) Possible context
RTX 5070 · 12 GiB 12.9 — does not fit
RTX 5070 Ti · 16 GiB 17.2 — missing 0.8 GB
RTX 4090 · 24 GiB 25.8 7.7 118k
RTX 5090 · 32 GiB 34.4 16.3 249k
MacBook Pro M4 Max · 48 GiB 51.5 33.5 262k (ceiling)
Mini PC · 96 GiB 103.1 85.1 262k (ceiling)
Arithmetic, not measurement: weights in Q4_K_M + KV cache at 64 KiB/token + 1 GB overhead. The weights are included without rounding (17.03 GB, and not the 17.0 displayed in the table above) — recalculating with 17.0 changes the remaining column to the first decimal place, and the context column does not change. Does not include the state of linear layers or what the system already occupies on the GPU. Measured numbers come in the next post.

Look at the second line. The 5070 Ti is out by 0.8 GB — less than 5% of the board’s memory. The weights alone occupy 17.03 of the 17.18 GB it has; there are 152 MB left for runtime and context, which isn’t even enough to load.

It’s the kind of margin that makes someone buy the wrong board, and the trap has two parts. The unit deceives in favor of the board: the “16 GB” on the box is 16 GiB, or 17.18 decimal GB — and suddenly a model “of 17 GB” seems to fit. The weights do fit. What doesn’t fit is what comes after them: the KV cache and runtime overhead are not optional, and they don’t appear on any label.

What you can reproduce now

The arithmetic part of this post you can check yourself without downloading 17 GB:

# The official configuration — the source of everything above
curl -sL https://huggingface.co/Qwen/Qwen3.6-27B/raw/main/config.json \
  | python3 -c "
import json,sys
from collections import Counter
t = json.load(sys.stdin)['text_config']
tipos = Counter(t['layer_types'])
plenas = tipos['full_attention']
kv = 2 * plenas * t['num_key_value_heads'] * t['head_dim'] * 2
print('camadas:', t['num_hidden_layers'], dict(tipos))
print('KV por token:', kv, 'bytes =', kv // 1024, 'KiB')
"

And to check on your own board after downloading:

ollama pull qwen3.6:27b

# With declared context — the runtime pattern is not what you will use
OLLAMA_CONTEXT_LENGTH=131072 ollama run qwen3.6:27b "ok"

# How much was actually reserved and if there was any left on the CPU
ollama ps
nvidia-smi --query-gpu=memory.used,memory.total --format=csv

What I still don’t know

Being explicit about the boundary between what this post proves and what it only suggests.

Proven here: the hybrid architecture, the KV cache of 64 KiB/token, and what the arithmetic says about each machine. All derived from the official config.json, reproducible by the command above.

Not yet measured: the state of linear layers, real GPU consumption after loading, tokens per second on each machine, and what linear attention costs in quality in a 200 thousand token context — because compressing the past into a fixed state is not free, and it’s exactly the test that matters.

That last one is the good question. A model that promises 262k of context and fits on a 4090 is news; if it remembers what was in token 30 thousand when it reaches 200 thousand is another story, and no one answers that with arithmetic.

In the next post I measure it, on six machines. The math says what fits; only measurement says what’s worth it.


Sources: official config.json on Hugging Face · model announcement · qwen3.6:27b on Ollama