Skip to content
Fantástico Mundo de Jon
RSS

2 posts

Posts tagged LLM local

All tags

  1. I swapped the local model in OpenCode: what sped up and what only looked like a model bug

    I was already coding with OpenCode and a Qwen on this GPU. I swapped 3.6 for 3.8. Here is what got faster, what hit more often, and the day I thought the model had gotten worse — it was configuration.

  2. When the model doesn't fit: the real cost of offloading to CPU

    In the previous post, I estimated that sending layers to the CPU costs ten times the performance. I tested on two GPUs and was wrong: the same model drops from 66 to 3.1 tokens per second. And nothing in the runtime warns when this happens.