Skip to content
Fantástico Mundo de Jon
RSS

4 posts

Posts tagged VRAM

All tags

  1. When the model doesn't fit: the real cost of offloading to CPU

    In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.

  2. Swapping the RTX 5060 Ti for the 5070 Ti doubles bandwidth, not capacity

    Both have 16 GB. The same model, quantization, and context fit in both. What doubles is the memory bandwidth, from 448 to 896 GB/s, and it's this that determines tokens per second.

  3. Qwen3.6 27B: how much VRAM does it really need

    A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.

  4. How much VRAM an LLM really uses

    The back-of-the-envelope math — parameters times bits — is off by several gigabytes, because it ignores the KV cache. Here is the full calculation, with the formula, the numbers and how to verify it on your own GPU.