Skip to content
Fantástico Mundo de Jon
RSS

6 posts

Posts tagged Local LLM

All tags

  1. The MacBook Pro M4 has more RAM than the 5090 and I stay with CUDA

    The 48 GB unified memory in the M4 holds more than the 5090's 32 GB. Nevertheless, vLLM, llama.cpp, and diffusion remain on NVIDIA because decoding is bandwidth and the stack I already use is CUDA.

  2. I ran Microsoft's BitNet on two CPUs: the 6x shrinks against the Q4 already on disk

    I cloned Microsoft's bitnet.cpp. The README promises 6.17x on CPU. On two machines at home the 4.2x versus f16 showed up; versus the Q4 Ollama already ships, it drops to 1.3x. And the model ships with the wrong activation.

  3. Swapping the RTX 5060 Ti for the 5070 Ti doubles bandwidth, not capacity

    Both have 16 GB. The same model, quantization, and context fit in both. What doubles is the memory bandwidth, from 448 to 896 GB/s, and it's this that determines tokens per second.

  4. Qwen3.6 27B: how much VRAM does it really need

    A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.

  5. How much VRAM an LLM really uses

    The back-of-the-envelope math — parameters times bits — is off by several gigabytes, because it ignores the KV cache. Here is the full calculation, with the formula, the numbers and how to verify it on your own GPU.

  6. Writing the edge case boosts accuracy from 46% to 90%

    360 generations, three local models, three ways of the same request. The big leap is between not writing the edge case and writing it: 46% to 90%. The markdown table, against the same running text, did not pay off what I expected.