1 post
Posts tagged offload
-
When the model doesn't fit: the real cost of offloading to CPU
In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.