2 posts
Posts tagged quantization
-
I ran Microsoft's BitNet on two CPUs: the 6x shrinks against the Q4 already on disk
I cloned Microsoft's bitnet.cpp. The README promises 6.17x on CPU. On two machines at home the 4.2x versus f16 showed up; versus the Q4 Ollama already ships, it drops to 1.3x. And the model ships with the wrong activation.
-
How much VRAM an LLM really uses
The back-of-the-envelope math — parameters times bits — is off by several gigabytes, because it ignores the KV cache. Here is the full calculation, with the formula, the numbers and how to verify it on your own GPU.