I ran Microsoft's BitNet on two CPUs: the 6x shrinks against the Q4 already on disk
I cloned Microsoft's bitnet.cpp. The README promises 6.17x on CPU. On two machines at home the 4.2x versus f16 showed up; versus the Q4 Ollama already ships, it drops to 1.3x. And the model ships with the wrong activation.
In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.
en Machine translation by grok-4.6, reviewed by the author. Read the Portuguese original
If you have ever asked ChatGPT, Grok or Claude Code for a snippet of code, the bill is the same. The model answers fast because it sits on a GPU that is not yours, in a building that is not yours. The text leaves the house and comes back ready.
Some people are trying to flip that bill: run the model on the processor on the desk, no graphics card, no file going to the United States. Microsoft has a project for that. It is called bitnet.cpp. Forty thousand stars on GitHub, a paper on arXiv, and a number at the top of the README: up to 6.17x faster on x86 CPUs.
Six times. Not twenty percent. Six.
The question almost nobody asks, and that the paper itself answers if you open section 4.1.2, is: six times faster than what?
I cloned it, compiled it and measured on two machines in this house. The 6.17x did not show up. What showed up was a 4.2x against an opponent nobody uses, a 1.3x against the file already on your disk, and a flagship model running with the wrong activation function from the day I downloaded the repository.
First, what each piece is
Five names that show up for the rest of the post, with no mystery.
BitNet is the idea. The model’s “weights”, the numbers it learned, stop being fractional and become three values: -1, 0 and +1. With no fractional numbers, multiplying turns into adding and subtracting. The energy arithmetic looks pretty. The conclusion people draw is that a big model now fits on a modest machine.
bitnet.cpp is Microsoft’s program that puts that idea on the CPU. That is what the README is selling.
llama.cpp is the engine almost everyone already has. Under the hood, Ollama is llama.cpp. If you run a local model in this house or in yours, it goes through here.
Q4_K_M is the compression Ollama hands you by default. Four bits per weight, give or take. It is the file already on the disk of anyone running models locally.
F16 is the model with no compression. Five gigs for a 2-billion-parameter model. Nobody runs that on CPU. It is the opponent Microsoft picked for the 6.17x.
There are two clocks, and people mix them up.
Prefill is the model reading your prompt. Decode is it writing the answer, one little piece at a time. The famous tok/s from hardware YouTube is almost always decode. Chat is almost all decode. RAG, the trick of pasting a document and asking a question, is almost all prefill. The number that decides whether cloning bitnet.cpp is worth it is the other one.
The 6.17x was measured against an opponent nobody uses
Before I show what I measured, it is worth opening Microsoft’s number. The answer sits in their own paper.
bitnet.cpp compares itself against llama.cpp in Float16. It is spelled out in section 4.1.2 of arXiv 2502.11880: “we chose llama.cpp Float16 as the baseline”.
So the baseline is a 2-billion-parameter model at full precision, taking up 5 GB, running on CPU. Nobody does that. People run Q4_K_M. The question that matters is not “how much does it gain over f16”. It is “how much does it gain over the file already on disk”.
The 1-bit file is not 1 bit
The first surprise came before any stopwatch, just looking at the file size.
The official GGUF, the package you download to run it, is 1,187,801,280 bytes for 2.41 billion parameters. That works out to 3.94 bits per weight, not 1.58. So I opened the tensors, the blocks of numbers inside the file, to find out where the money went.
Anatomy of ggml-model-i2_s.gguf
| Part | Parameters | Bytes | Bits per weight |
|---|---|---|---|
| Ternary weights (I2_S) | 2,084,044,800 | 521,017,920 | 2.000 |
| Embedding table (F16) | 328,335,360 | 656,670,720 | 16.000 |
| Normalizations (F32) | 440,320 | 1,761,280 | 32.000 |
| Whole file | 2,412,820,480 | 1,187,801,280 | 3.938 |
Two things come out of that, and both contradict the pitch.
The ternary weight, the one with three values, costs exactly 2 bits, not 1.58. The 1.58 is log₂(3), the theoretical entropy of three symbols. Getting close to it would require packing five values into eight bits, which is what llama.cpp’s TQ1_0 does. Microsoft’s I2_S format does not: it stores four weights per byte, two bits each, and leaves the fourth state unused.
And there is a huge table at the front of the model, the vocabulary. Every possible word becomes a vector. That table is not ternary. It eats 55.3% of the file by itself. The model uses the Llama 3 tokenizer, with 128,256 vocabulary entries. In F16 that is 656 MB the 1-bit technique never touched. The smaller the ternary model, the bigger the slice the vocabulary takes, and that is why the promised saving does not scale downward.
Where I measured
Two machines, deliberately with different vector instructions. One is the desktop I work on. The other is a mini PC that serves as an infrastructure server here at home and has no GPU at all, which is exactly the scenario BitNet was designed for.
The two machines
- titan
- Intel i7-12700 · 8P+4E, 20 threads · AVX2 + AVX-VNNI, no AVX-512 · 31 GiB · Ubuntu 24.04
- aegis
- AMD Ryzen 7 7840HS · 8 cores, 16 threads · full AVX-512 with VNNI · 58 GiB · openSUSE Leap 16
- bitnet.cpp
- commit 0b341e5, 2026-07-27 · llama.cpp submodule 390c3077
- Compiler
- clang 22.1.8 · cmake 4.4.2 (via micromamba, no root)
- Model
- microsoft/BitNet-b1.58-2B-4T-gguf · ggml-model-i2_s.gguf
- Method
- llama-bench -p 128 -n 32 -r 3, sweeping 4/8/12/16/20 threads
One methodological detail that changes the result: do not use the repository’s e2e_benchmark.py to measure prefill. It calls llama-bench with -b 1, and with batch 1 the 512-token prompt goes in one token at a time. Prefill stops being matrix-by-matrix multiplication and becomes matrix-by-vector, which drags the number down to decode speed. I called llama-bench directly, with the default batch.
Every format below came from the same checkpoint. I converted the official bf16 weights to GGUF f16 and quantized from there. Architecture, tokenizer and weight values are identical across all rows. The only thing that changes is the kernel, the piece of code doing the arithmetic.
The gain depends on who you compare against
titan · Intel i7-12700
| Format | File | Decode tok/s | Prefill tok/s |
|---|---|---|---|
| I2_S (bitnet.cpp) | 1.10 GiB | 49.2 | 345.8 |
| TQ2_0 (llama.cpp) | 0.75 GiB | 68.0 | 207.3 |
| TQ1_0 (llama.cpp) | 0.66 GiB | 49.1 | 87.5 |
| Q4_K_M | 1.41 GiB | 32.4 | 145.8 |
| Q8_0 | 2.39 GiB | 21.7 | 82.0 |
| F16 | 5.11 GiB | 11.6 | 101.1 |
aegis · AMD Ryzen 7 7840HS
| Format | File | Decode tok/s | Prefill tok/s |
|---|---|---|---|
| I2_S (bitnet.cpp) | 1.10 GiB | 43.8 | 580.1 |
| TQ2_0 (llama.cpp) | 0.75 GiB | 64.0 | 244.7 |
| TQ1_0 (llama.cpp) | 0.66 GiB | 65.2 | 118.9 |
| Q4_K_M | 1.41 GiB | 34.9 | 259.8 |
| Q8_0 | 2.39 GiB | 21.4 | 180.7 |
| F16 | 5.11 GiB | 10.4 | 177.6 |
Look at the table and ask: I2_S beat whom?
How much I2_S wins by, on decode, depending on the opponent
| Comparison | titan | aegis |
|---|---|---|
| I2_S over F16 — the README number | 4.24x | 4.20x |
| I2_S over Q4_K_M — what you use today | 1.52x | 1.26x |
| TQ2_0 over I2_S — llama.cpp against the official | 1.38x | 1.46x |
The 4.2x against f16 is real and showed up practically the same on both machines. I did not reach 6.17x, and I did not expect to: that number came from synthetic, untrained models on a 13th-generation i7, and the paper itself says so in the fine print.
Look at the second row. Against Q4_K_M, the file already on your disk, the gain drops to somewhere between 1.26x and 1.52x. That is good. It lands in adjustment territory, nowhere near a change of category. And it comes with a price no benchmark shows, which I get to below.
The llama.cpp you already have wins at chat
When TQ2_0 beat I2_S, I redid the measurement to be sure.
Mainline llama.cpp has had its own ternary types since 2024, TQ1_0 and TQ2_0, which have nothing to do with bitnet.cpp. I quantized the same checkpoint to TQ2_0 purely out of curiosity, and it decoded 1.38x faster than I2_S on titan and 1.46x faster on aegis. In a file 32% smaller, because TQ2_0 also compresses that vocabulary table that I2_S leaves in F16.
bitnet.cpp does not lose across the board. On prefill it dominates, and by a wide margin: 345 against 207 tok/s on titan, 580 against 244 on aegis. That is its kernel chewing through the prompt, which is where matrix-by-matrix multiplication appears and where the ternary trick pays off most.
In plain language: Microsoft’s kernel is better at swallowing the prompt, llama.cpp’s is better at spitting out tokens. If your workload is RAG, classification, anything that reads a lot and writes little, bitnet.cpp pays off. If it is chat, plain llama.cpp already serves you better, and you do not need to clone anything.
It is worth noting where the mini PC beat the desktop. The Ryzen’s 580 tok/s of prefill against the Intel’s 345 is not about clock or core count. It is AVX-512 with VNNI, a way for the processor to do the same sum on several numbers at once. The i7-12700 lacks it because Intel disabled it across the whole 12th generation. For ternary workloads on CPU, that instruction is worth more than the brand on the processor.
The model ships with the wrong function
This is where the benchmark turned into something else.
An activation function is the seasoning in the middle of the network. The model learned with one. The code applies another. Speed does not even change, because the two functions cost the same to the processor. Quality changes, and it changes badly.
The config.json of BitNet b1.58 2B declares "hidden_act": "relu2". The model was trained with squared ReLU in the feed-forward, an unusual and deliberate choice. Except the llama.cpp graph embedded in bitnet.cpp, in the file src/models/bitnet.cpp, line 133, does this:
grep -n 'LLM_FFN_SILU' 3rdparty/llama.cpp/src/models/bitnet.cpp 133: LLM_FFN_SILU, LLM_FFN_PAR, il);
SiLU, not relu². The graph is applying a function the model did not learn.
I did not want to take a code reading on faith, so I measured. I ran perplexity on wikitext-2. Perplexity is a way to measure whether the model is “surprised” by the text: the lower the number, the more the text looks like something it was trained on. 17 is Microsoft’s number. I ran with the repository as it ships, swapped LLM_FFN_SILU for LLM_FFN_RELU_SQR, recompiled, and ran again. Same model, same text file, same 100 chunks, same machine.
Perplexity on wikitext-2, one line of difference
| Build | Perplexity | Deviation |
|---|---|---|
| As the repository ships today (SiLU) | 98.80 | ± 2.26 |
| Swapping one line to relu² | 17.85 | ± 0.34 |
| Reference published by Microsoft | 17.11 | ± 0.13 |
The model comes out 5.5 times worse than it should. And the corrected value matches Microsoft’s own reference, which confirms that the only thing wrong was indeed that line.
Except that one-line swap works for my benchmark and does not work as a real fix. The class llama_model_bitnet_b158, which is Microsoft’s model, inherits the graph from llama_model_bitnet, which is the old 1bitLLM architecture. And that one genuinely uses SiLU. Swapping the line fixes the new model and breaks the old one, which then runs with the wrong activation in the opposite direction.
A fix that could go upstream has to tell the two apart: either b158 gets its own graph, or the activation starts being read from the GGUF metadata. That is the design a llama.cpp maintainer already asked for in one of the closed pull requests. Until someone does it, whoever clones the repo picks which of the two models will run wrong.
After measuring, I went digging. The problem is already known: it has been filed as issue #602 since August 8, and there are two pull requests fixing it in mainline llama.cpp. Both were closed by their own authors, without being rejected by anyone. The bug is still there.
The documented path to generate f16 does not run
To build the tables above I needed the same model in f16. The repository documents how to do that. The converter accepts --outtype f16. On today’s clone, that does not finish.
The file 3rdparty/llama.cpp/gguf-py/gguf/constants.py has two dictionaries, one that names model architectures and another that names vision projectors. The two lines registering BitNet were pasted into the second one:
# 3rdparty/llama.cpp/gguf-py/gguf/constants.py, lines 1131-1143
VISION_PROJECTOR_TYPE_NAMES: dict[VISION_PROJECTOR_TYPE, str] = {
VISION_PROJECTOR_TYPE.MLP: "mlp",
# ...
VISION_PROJECTOR_TYPE.STEP3VL: "step3vl",
MODEL_ARCH.BITNET: "bitnet", # <- these two
MODEL_ARCH.BITNET_25: "bitnet-25", # <- are in the wrong dictionary
}
The result is that MODEL_ARCH.BITNET_25 is the only architecture in the entire project with no name in the right dictionary, and the converter dies with KeyError before writing the first tensor. I fixed it by moving the line to the correct dictionary, and then found the second layer: the name "bitnet-25" does not exist in the C++ runtime, which only knows "bitnet" and "bitnet-b1.58". The converter was producing a file the project itself cannot open.
I verified that the 332 generated tensors are identical to the official GGUF’s and pointed BITNET_25 at "bitnet-b1.58". Then the converter closed the file and the runtime opened it.
The practical effect: the documented path to generate the f16 baseline is broken. The speed comparison the README uses to claim 6.17x is not reproducible by anyone cloning the repository today. The only artifact that works is the pre-baked GGUF that Microsoft publishes.
If you want to redo this
I left all three fixes and the whole walkthrough in two forks, because a patch you cannot apply is worth nothing:
- jordaobass/llama.cpp — all three patches: the activation, the architecture registration, and the macOS
src1_cont - jordaobass/BitNet — the submodule already pointing there, plus
BANCADA.mdwith the step by step and the rawllama-benchJSON from both machines
Cloning with --recursive and following BANCADA.md should reproduce the tables above, including the f16 baseline step that breaks on the official repository. Nothing there needs root: cmake and clang come via micromamba, because neither of my two machines had a compiler installed.
What the 1.3x costs
I still have to say what you pay for the 1.3x.
BitNet b1.58 2B has 4,096 tokens of context. Context is short-term memory: the file it read, the question, the previous answer. It is in the config.json, it is not a runtime limitation. ChatGPT, Grok and Claude Code work with tens or hundreds of thousands. In 2026, 4096 rules the model out of practically anything agentic, any serious RAG, or reading a document.
In Microsoft’s own quality comparison, the model averages 54.19 against 55.23 for Qwen2.5-1.5B. It loses, in the table made by the people who built it. It wins on GSM8K and common-sense reasoning, and loses badly on MMLU and on code.
And that is the ceiling of the ecosystem. There is no natively pretrained ternary model above 3B: the largest is Falcon-E-3B, and Microsoft never released the 7B or 70B weights that appear in the papers. Everything above that is conversion after training, and the two best recipes, Intel’s and Prism ML’s, have not published their conversion code.
I tried it on a Mac and never got to the number
The third architecture was missing, so I freed up disk on the MacBook (M4 Pro, 10P+4E, macOS 26.5) and went to run the same benchmark. I never got to the tokens per second. The road there gave me more than the number would have.
The default build does not pass on macOS. Two stops, in sequence. The first is the linker refusing libggml-base.dylib, because it references quantize_i2_s and dequantize_row_i2_s, which are defined in ggml-cpu. On Linux, ELF tolerates undefined symbols in a shared library and nobody notices; macOS Mach-O does not. It passes with -Wl,-undefined,dynamic_lookup.
The second is code that does not compile. The I2_S block in ggml-cpu.c reads src1_cont, which is declared inside a #if GGML_USE_LLAMAFILE on line 1361, and the I2_S block sits outside any guard:
clang -c ggml-cpu.c ggml-cpu.c:1495:40: error: use of undeclared identifier ‘src1_cont’
Anywhere that macro does not get defined, the project simply does not build. It is a one-line fix, and it became the third patch in the fork.
The kernel Microsoft optimized for ARM does not compile either. This is the strangest part, because TL1 is the official path on ARM and the demo video in the README itself is an Apple M2. Turning on BITNET_ARM_TL1=ON, bitnet-lut-kernels.h asks for tensor->backend and GGML_BACKEND_TYPE_CPU, two things ggml removed from the API. The kernel generator fell behind the very llama.cpp the project carries as a submodule.
And with Metal in the build, I2_S dies quietly. In the default macOS configuration, which includes Metal, loading the official GGUF takes the process down with SIGSEGV and no message at all. Turning Metal and Accelerate off, it starts running.
That is where I stop, for a boring reason. The machine was in no state to be measured: a Virtualization.framework VM was eating two cores, several iOS simulators were booted, and the load average was 387. The prefill that came out of it, 9.26 tok/s, does not let me separate a scalar fallback from CPU contention. A number I cannot read does not go into a table. That waits until I have the machine quiet.
I do not have a tok/s number on ARM to publish. What I have is the state of the code in August 2026: bitnet.cpp does not build on macOS without two workarounds, and the kernel it advertises for ARM does not build at all.
What I would do
If you have a GPU, none of this matters and the post ends here. ChatGPT, Grok and Claude Code stay in the cloud; the Qwen on this card stays on the card. A 2B BitNet with 4096 of context does not enter that conversation.
If you want to run on CPU, I would not clone bitnet.cpp. I would take the llama.cpp you already have, quantize to TQ2_0 and get on with it. It decodes faster, the file is smaller, it runs on more backends, and the activation comes out right. Compiling bitnet.cpp is only worth it if your workload is almost all prefill, and then it is worth quite a lot.
Before believing 6.17x, ask against what. If the answer is f16 on CPU, redo the arithmetic with the Q4_K_M sitting on disk.