2 posts
Posts tagged llama.cpp
-
The MacBook Pro M4 has more RAM than the 5090 and I stay with CUDA
The 48 GB unified memory in the M4 holds more than the 5090's 32 GB. Nevertheless, vLLM, llama.cpp, and diffusion remain on NVIDIA because decoding is bandwidth and the stack I already use is CUDA.
-
I ran Microsoft's BitNet on two CPUs: the 6x shrinks against the Q4 already on disk
I cloned Microsoft's bitnet.cpp. The README promises 6.17x on CPU. On two machines at home the 4.2x versus f16 showed up; versus the Q4 Ollama already ships, it drops to 1.3x. And the model ships with the wrong activation.