Skip to content
Fantástico Mundo de Jon
RSS

I run AI models on my own hardware and publish the numbers.

Six machines, from 12 to 96 GB: four NVIDIA GPUs, Apple Silicon and unified memory. Lab

Writing

  1. The MacBook Pro M4 has more RAM than the 5090 and I stay with CUDA

    The 48 GB unified memory in the M4 holds more than the 5090's 32 GB. Nevertheless, vLLM, llama.cpp, and diffusion remain on NVIDIA because decoding is bandwidth and the stack I already use is CUDA.

  2. I swapped the local model in OpenCode: what sped up and what only looked like a model bug

    I was already coding with OpenCode and a Qwen on this GPU. I swapped 3.6 for 3.8. Here is what got faster, what hit more often, and the day I thought the model had gotten worse — it was configuration.

  3. I ran Microsoft's BitNet on two CPUs: the 6x shrinks against the Q4 already on disk

    I cloned Microsoft's bitnet.cpp. The README promises 6.17x on CPU. On two machines at home the 4.2x versus f16 showed up; versus the Q4 Ollama already ships, it drops to 1.3x. And the model ships with the wrong activation.

  4. When the model doesn't fit: the real cost of offloading to CPU

    In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.

  5. Swapping the RTX 5060 Ti for the 5070 Ti doubles bandwidth, not capacity

    Both have 16 GB. The same model, quantization, and context fit in both. What doubles is the memory bandwidth, from 448 to 896 GB/s, and it's this that determines tokens per second.

  6. Coloring Pages with Local AI: From Prompt to a File That Prints Well

    Generating the image is the easy part, and that's where most tutorials end. What determines whether the child can paint is what comes next: binarizing, vectorizing, and printing without jagged edges.

  7. Qwen3.6 27B: how much VRAM does it really need

    A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.

  8. How much VRAM an LLM really uses

    The back-of-the-envelope math — parameters times bits — is off by several gigabytes, because it ignores the KV cache. Here is the full calculation, with the formula, the numbers and how to verify it on your own GPU.

See all posts