Skip to content
Fantástico Mundo de Jon
RSS

I swapped the local model in OpenCode: what sped up and what only looked like a model bug

I was already coding with OpenCode and a Qwen on this GPU. I swapped 3.6 for 3.8. Here is what got faster, what hit more often, and the day I thought the model had gotten worse — it was configuration.

Updated

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by grok-4.6, reviewed by the author. Read the Portuguese original

If you use ChatGPT to write code, you know the scene. You ask for a filter, a button, an endpoint. Something that looks right comes back. Sometimes it is. Sometimes it compiles and is wrong in the detail you did not read.

I do the same thing, except the “brain” is not in the cloud. It is a graphics card on this desk, an RTX 5090. The program that drives the model is OpenCode: it reads files, edits, runs tests. An intern who does not get tired and does not send the repository out of the house.

The model I was using was Qwen3.6, 27 billion parameters. “27B” is the size of the head. Bigger, in theory, means it knows more — and it eats more video memory. Then Qwen shipped 3.8. Same size class, same free license (Apache 2.0), and now it can look at images.

The question everyone asks, including me: did it get faster? Does it hit more often?

I will tell you what I measured on this desk. And the day I thought the model had gotten worse after the swap — and it was me who had configured it wrong.

First, what each piece is

Three names that show up for the rest of the post, with no mystery.

OpenCode is the app. It opens the project, calls tools, writes the git diff. Without it, the model only chats.

Qwen is the model. 3.6 is what I had. 3.8 is the successor. Both fit on this card. Both speak enough Portuguese and English for my day job in .NET, JavaScript and Flutter.

Ollama is the engine that loads the model on the GPU and answers on localhost. OpenCode talks to it. If Ollama advertises a context the card cannot deliver, the agent “gets dumb” and it looks like Qwen’s fault.

Local, here, means: the code does not go to OpenAI’s API or Anthropic’s. The GPU works, the disk is mine, the test runs in the repo folder.

Two clocks, and people mix them up

There is the GPU clock. How many little pieces of text the model writes per second, the famous tok/s. I measure that with curl and nvidia-smi. It is the number hardware YouTube loves.

There is the task clock. From “implement the filter” until dotnet test — or vitest, or flutter test — passes. That clock includes the model describing the edit instead of editing, inner reasoning eating the answer, the program announcing 262 thousand tokens of memory and throwing the instructions away in the middle of a file. Then I open another session and lose ten minutes.

On 3.6 the two clocks already fought. I put OpenCode on three little coding tasks (a parity function, an LRU cache, an interval merge). Qwen3.6 in Q6 — Q6 is a compression of the model, a bit heavier and a bit more faithful — scored 3/3. It tied Haiku 4.5, a paid Anthropic model, on accuracy. Haiku finished in about 24 seconds. The local one, between 56 and 103.

Same hits. Two to four times the wall clock.

In plain language: if your metric is beating ChatGPT on a stopwatch, local on this card still loses. If the metric is the patch compiled, the test passed, and the repo did not fly to the United States, 3.6 was already usable. 3.8 sits on top of that.

What was on the desk · August 2026

GPU
RTX 5090 · 32 GB of video memory
Engine
Ollama 0.32.15 (llama.cpp underneath)
App
OpenCode 1.18.25, talking straight to Ollama on port 11434
Now
Qwen3.8 27B Q6 with vision · 131 thousand tokens of context
Before
Qwen3.6 27B Q6 with vision, same OpenCode

What the card actually gained in speed

The architecture of both is the same. 64 floors in the building, only 16 with that “scratch notebook” that grows as the conversation goes — the KV cache. The 3.6 post did not age.

What changed in the package was a vision projector (the model can look at a screenshot) and, when it writes, more speed.

Write speed on the 5090 · no inner reasoning · short context

Model Video memory Writes (tok/s) Reads the prompt (tok/s)
Qwen3.6 27B Q4 16.95 GB 76.0 1372
Qwen3.8 27B Q4 18.25 GB 111.8 1141
RTX 5090, Ollama 0.32.15, same technical prompt, 256 output tokens, temperature 0, thinking off, context 8,192. Median of 3 runs after a warmup. Detail in the 3.6 → 3.8 swap post, 21 Aug 2026. I treat 3.8 as around 100 to 110 tok/s, not as an exact 1.5×.

Around 100 to 110 tokens per second against 76. You feel that in the stretch where the model is typing the file.

You barely feel it in the stretch where it reads the file. In this battery 3.8 read slower: 1141 vs 1372. And a coding agent reads much more than it writes. It opens the service, opens the test, opens the component, and only then changes three lines.

So: buying the swap only because it “writes faster” is watching the wrong clock.

The other number I measured, already on 3.8 compressed as Q6 and with vision, was this. Minimal request: “use the tool”. The model returned a real command, not a paragraph saying “I will edit the file”. 14 seconds, 67 tokens. On 3.6 that flickered. When it narrates the tool, OpenCode does not edit. The task is not slow. It simply does not happen.

Qwen’s brochure says it hits more often on code

I did not run this. These are numbers Qwen put on the 3.8 card, comparing against 3.6.

Think of these names as standardized exams. SWE-bench is “can it fix a bug in a real repository?”. Terminal Bench is “can it find its way around a terminal?”. LiveCodeBench is “can it pass a live programming exercise?”.

Qwen’s numbers, not this desk

Exam Qwen3.6 27B Qwen3.8 27B
Terminal Bench 2.1 63.4 73.0
SWE-bench Pro 53.5 61.7
DeepSWE 1.1 13.3 42.2
LiveCodeBench v6 83.9 90.3
Values from the Qwen3.8-27B model card, 14 Aug 2026. DeepSWE and SWE-bench Pro use a Claude Code-style harness, 256k context. Not my measurement. Artificial Analysis, which scores models on its own, published Intelligence Index 38 → 52 on the same comparison.

The jump they declare is in “use a tool and touch a project”, not science questions. If you swap hoping the model aces a quiz, the table does not promise that. If you swap because 3.6 was already your local 27B coder, that is exactly the case they point at.

On this desk, 3.6 Q6 had already tied Haiku on the three little tasks. The more squeezed version, Q4, scored 1/3 on the same test. Compression moved that accuracy more than a family swap would move a parity function. That is why I stay on Q6 to drive tools, and leave Q4 for chat. I did not go back.

I blamed the model. It was configuration.

After the swap, 3.8 “hallucinated” and stopped mid-sentence. Classic blame: the new model is worse. I went and read the log of the server Ollama starts underneath.

Three defects. None of them the weights.

First: the box says 262 thousand tokens of context. Context is the conversation’s short-term memory — the file it read, the test it ran, the instructions. 262 thousand is what the model accepts. On this card, with Q6, 262 thousand pushes a slice of the model into ordinary RAM, off the GPU. Reading the prompt drops 29%. And there is an ugly detail: the server tries to “slide” memory when it fills up, the hybrid architecture will not let it, and the program keeps only 4 tokens. The instructions are gone. The model starts answering a question nobody asked. It looks like hallucination. It is configured amnesia.

The ceiling that fits entirely on the GPU, the way Ollama is set today, is about 139 thousand. I locked OpenCode at 131 thousand. 262 thousand is the number on the box. It is not the number of this 5090.

Second: the answer limit too short. 3.8 thinks out loud before it answers. If you leave only 512 tokens of ceiling, the thought eats everything and the answer cuts on the first item. At 2048, all five items come out. Measured on this card, same prompt, only the ceiling changed.

Third: OpenCode going through a middleman. I had a program in the middle, LiteLLM. The text stream came in a format the old OpenCode rejected. The path that works here is OpenCode talking straight to Ollama, on port 11434. The middleman stays for wiki, project keys, and cloud models.

When the agent “got dumb” after the swap, most of the time it did not. I had let the engine lie about memory.

How this shows up in .NET, JavaScript and Flutter

How I ask is the same in all three. I open OpenCode in the project folder. The model is the one on this desk. It may edit and it may run a command. I do not ask “write the system”. I ask for the task with the contract in view: what goes in, what comes out, what the test has to break.

In .NET, a feature folder fits in those 131 thousand tokens with the test. House rules — Result, ProblemDetails, no exceptions for control flow — live in the repo, and the local agent reads them. The gain I feel versus 3.6 is not a faster dotnet build. It is the model calling read-file and edit-file, instead of pasting a snippet in the chat. When it narrates, I lose the whole turn. It is like an intern describing the fix on Slack and never opening Visual Studio.

In JavaScript I have a test number: the three OpenCode tasks on 3.6 Q6 passed. Day to day it is Next, a component, a test. 3.8 with vision helps when the screen is wrong. The screenshot goes into the conversation, the component file too. Before that I described the layout in prose (“the button ended up on the left, with a weird gap”) and it invented flex that was not in the pixel.

Flutter is the same screen logic. Emulator screenshot plus the widget. House rules stay spec, not a miracle of the model. 3.8 did not spare me from writing the weird case. I already measured that: writing what happens in the odd case lifts accuracy from 46% to 90%. A pretty markdown table, against the same running prose, barely pays.

The time I gained, in all three, is less rework. A session that does not die in inner thought. A tool that fires. Memory that does not drop the instructions in the middle of a file.

Writing at 100 tokens per second is the bonus that shows up when it is already touching the right file.

What I would do in your place

If you already use OpenCode with 3.6 27B Q6 on a 24 GB or larger card, I would swap. It costs about 1.3 GB extra on the GPU and a recent Ollama (0.32.12 and up). Writing goes up. Qwen’s brochure promises the gain where you use it — touching a project, not acing a quiz. Lock context to what your card delivers. Here, 131 thousand. And point OpenCode straight at Ollama, no middleman.

If your comparison is Haiku in the cloud, set the clock. Local is still slower per small task and still wins on data that does not travel. Those are different ledgers. I use both. Local 3.8 got good enough for the turn where the repository cannot leave the building.

If the card has 16 GB, 3.8 does not fix what 3.6 already did not. The video-memory post still stands.

I did not run SWE-bench on this desk. This post proves write speed, a real tool command, the actual 131 thousand token ceiling, and what breaks when you believe the 262 thousand on the box. The rest is what Qwen published.

A vendor number is a starting point. It is not a closing argument. Run it on your card and watch the test pass — or not.