Skip to content
Fantástico Mundo de Jon
RSS

2 posts

Posts tagged benchmark

All tags

  1. When the model doesn't fit: the real cost of offloading to CPU

    In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.

  2. Writing the edge case boosts accuracy from 46% to 90%

    360 generations, three local models, three ways of the same request. The big leap is between not writing the edge case and writing it: 46% to 90%. The markdown table, against the same running text, did not pay off what I expected.