Skip to content
Fantástico Mundo de Jon
RSS

3 posts

Posts tagged GPU

All tags

  1. When the model doesn't fit: the real cost of offloading to CPU

    In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.

  2. Qwen3.6 27B: how much VRAM does it really need

    A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.

  3. How much VRAM an LLM really uses

    The back-of-the-envelope math — parameters times bits — is off by several gigabytes, because it ignores the KV cache. Here is the full calculation, with the formula, the numbers and how to verify it on your own GPU.