Skip to content
Fantástico Mundo de Jon
RSS

All posts

  1. The MacBook Pro M4 has more RAM than the 5090 and I stay with CUDA

    The 48 GB unified memory in the M4 holds more than the 5090's 32 GB. Nevertheless, vLLM, llama.cpp, and diffusion remain on NVIDIA because decoding is bandwidth and the stack I already use is CUDA.

  2. I swapped the local model in OpenCode: what sped up and what only looked like a model bug

    I was already coding with OpenCode and a Qwen on this GPU. I swapped 3.6 for 3.8. Here is what got faster, what hit more often, and the day I thought the model had gotten worse — it was configuration.

  3. I ran Microsoft's BitNet on two CPUs: the 6x shrinks against the Q4 already on disk

    I cloned Microsoft's bitnet.cpp. The README promises 6.17x on CPU. On two machines at home the 4.2x versus f16 showed up; versus the Q4 Ollama already ships, it drops to 1.3x. And the model ships with the wrong activation.

  4. When the model doesn't fit: the real cost of offloading to CPU

    In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.

  5. Swapping the RTX 5060 Ti for the 5070 Ti doubles bandwidth, not capacity

    Both have 16 GB. The same model, quantization, and context fit in both. What doubles is the memory bandwidth, from 448 to 896 GB/s, and it's this that determines tokens per second.

  6. Coloring Pages with Local AI: From Prompt to a File That Prints Well

    Generating the image is the easy part, and that's where most tutorials end. What determines whether the child can paint is what comes next: binarizing, vectorizing, and printing without jagged edges.

  7. Qwen3.6 27B: how much VRAM does it really need

    A 27B model that uses half the KV cache of a Llama 3 8B. The hybrid architecture breaks the usual calculation — and decides, for less than 1 GB, which GPUs are left out.

  8. How much VRAM an LLM really uses

    The back-of-the-envelope math — parameters times bits — is off by several gigabytes, because it ignores the KV cache. Here is the full calculation, with the formula, the numbers and how to verify it on your own GPU.

  9. Writing the edge case boosts accuracy from 46% to 90%

    360 generations, three local models, three ways of the same request. The big leap is between not writing the edge case and writing it: 46% to 90%. The markdown table, against the same running text, did not pay off what I expected.

  10. The PRD comes first: the spec that prevents the agent from writing the wrong code

    The model doesn't err due to stupidity: it errs because it closed the gap left by the prompt on its own, and did so in a plausible way. What changes in the code when the contract comes before the request.

  11. Dissecting an Ollama Modelfile to Stop the Model from Improvizing in Code

    Ollama patterns are for conversation. For code, the task is to reduce randomness in sampling, write a dry SYSTEM, and save it in a derived model. Step-by-step with Qwen2.5-Coder.

  12. How Much Does It Cost to Program with ChatGPT: $20 Plus vs. API

    The choice between subscription and API is wrong in both directions, for the same reason: in chat, each message resends the entire conversation. The cost doesn't grow with the number of questions; it grows with its square.

  13. A website from scratch via chat: HTML is ready, CSS isn't

    I built an entire website by copying and pasting from the chat. The structure was almost ready, but the presentation wasn't — because HTML has a correct answer and layout only has a correct answer for a context that the model doesn't know.

  14. A complete API through chat: what GPT-4o gets right and where it misleads you

    In November 2024, no agent edits my files: the API exits the chat in blocks that I paste by hand. GPT-4o gets the structure right and misleads me in four places — the worst is the test that passes without testing anything.