Writing the edge case boosts accuracy from 46% to 90%
360 generations, three local models, three ways of the same request. The big leap is between not writing the edge case and writing it: 46% to 90%. The markdown table, against the same running text, did not pay off what I expected.
In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.
en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original
In July, I argued that replacing prose prompts with spec files made the agent more accurate. The post is here: contract in table format, edge cases, out of scope items, acceptance criteria, and a results table with blank cells and the promise to fill them later.
In October, I filled it in. The thesis from July, as I had written it, didn’t hold up. The structure of the file wasn’t what moved the needle.
What made the difference was something else, and that stood out stronger than the original thesis. Between a one-sentence request and the same task with edge cases written somewhere, the accuracy in edge cases went from 46% to 90%. Almost double. Between the spec arranged in markdown and the same content thrown into a continuous paragraph, the difference was 3.6 points — small, stable across repetitions, and in favor of the continuous text.
In other words: the format in which you arrange the request matters little to the model. Having decided and written what happens in the strange case matters a lot. That’s what I take from this experiment, and that’s what changes what I do tomorrow morning when I open Aider.
The extent of the fall of the structural thesis, to be specific: in the run without randomness, the arm with the spec and the arm with exactly the same information in paragraph form gave identical results in 18 out of 30 cells. In the slugify task, with devstral:24b at temperature 0, both returned the same code. It wasn’t similar: it was byte by byte the same file.
Bench · October 2025
- Placa
- RTX 5090 · 32 GB
- Modelos
- qwen3-coder:30b · devstral:24b · qwen2.5-coder:32b
- Quantização
- Q4 via Ollama
- Acesso
- Raw API, one turn, no history
- Tarefas
- 10 small pure functions in JS
- Braços
- prose · long prose · spec
- Gerações
- 360 (90 + 270)
- VRAM occupied
- 20 · 17 · 24 GB, in the order of the models
- Time for the two runs
- 15.2 min total call time
Why there are three arms, not two
Comparing “one sentence” with a “60-line spec” doesn’t answer anything. Of course, more information helps. That’s obviousness with the face of a result, and it was the mistake I most wanted to avoid.
If I only compared short prompt with spec, I wouldn’t know how to separate two different things: the fact that the information exists and the fact that it is in a table with headers. The first is content. The second is form. In July, I was selling the second. I needed to measure both.
That’s why the bench has three arms, and the middle one gives meaning to the other two:
- prose — one-sentence request, only happy path, no edge cases mentioned.
- long prose — exactly the same information as the spec, in continuous text, without structure.
- spec — the same information in structured markdown, with the four usual blocks.
Thus, prose vs. long prose measures information, and long prose vs. spec measures form. The second comparison was July’s thesis. It didn’t hold up. The first is the finding that remained, and it justifies the rest of this post.
The rest of the design exists to close easy exits. Each task’s tests are separated into happy path and edge cases, and each edge case corresponds to an item that the spec declares and the short prose omits, one-to-one. The output format instruction is the same in all three arms; otherwise, I would be measuring format along with content. Each task brings a reference implementation that needs to pass 100% of its own tests before the run starts: if the reference fails, the wrong thing is the test, not the model. And the generated code runs in an isolated process with a 10s timeout because model code doesn’t deserve trust. It enters a loop, calls process.exit, consumes memory.
The entire harness is in the repository. scripts/bancada-spec.mjs runs the race, scripts/bancada-spec/tarefas.mjs stores the ten tasks with the three prompts and tests for each, and scripts/bancada-spec/executor.mjs runs a candidate in its own process.
pnpm bancada --modelos qwen3-coder:30b,devstral:24b,qwen2.5-coder:32b --amostras 3 --temperatura 0.7 referencias validadas: 10 tarefas, todas em 100% bancada: 3 modelos × 10 tarefas × 3 braços, temperatura 0.7
The --retomar takes advantage of what has already run. You will need it because 270 generations don’t fit into a session without something falling in the middle.
First run: temperature 0, 90 generations
One request per cell, with no randomness at all. This is the run that answers “with this exact prompt, what comes out”.
| Arm | Happy Path | Edge Cases | % Edge | 100% Tasks |
|---|---|---|---|---|
| prose | 78/116 | 93/200 | 47% | 0/30 |
| long prose | 116/120 | 185/207 | 89% | 18/30 |
| spec | 110/120 | 183/207 | 88% | 14/30 |
The first column is what I already expected, and it’s not news: the model almost always gets the happy path right, in any arm. Everyone understands “format a number in Brazilian reais”.
The edge case column is where the real story happens. 47% against 89%. It’s the same phenomenon I described in July, now with numbers on top. The model doesn’t make mistakes because it’s stupid. It makes mistakes because it closed the gap that the request left open by itself. When no one says what to do with null, with negative values, with an empty string, it invents a plausible answer — and plausible fails the test.
And then comes the third line, which is where my July thesis dies without drama. The structured spec was one point below the same information thrown into a paragraph. A one-point difference I don’t claim. But I needed the spec to be ahead, and it wasn’t.
Read again the comparison that matters: short prose against any of the two informed arms. Almost double the accuracy in edge cases. That’s what pays for the cost of sitting down and writing the behavior before asking for the code.
Second run: 270 generations at temperature 0.7
A deterministic run doesn’t distinguish real difference from luck. I repeated three times at 0.7, which is where these tools usually run in real life.
| Arm | Edge Cases | % Edge | Per Repetition |
|---|---|---|---|
| prose | 277/607 | 46% | 46%, 48%, 43% |
| long prose | 554/614 | 90% | 88.0%, 91.3%, 91.3% |
| spec | 538/621 | 87% | 87.4%, 86.5%, 86.0% |
Dispersion matters more than the average here. Short prose oscillates five points between repetitions, while spec oscillates just over one. And the three spec repetitions were below the worst long prose repetition, meaning that the 3.6 point difference didn’t come from a bad round. It also doesn’t make a verdict, and the note below explains why.
What repeats, stable, in the three rounds, is the abyss between not writing the edge case and writing it. 46% on one side. 90% and 87% on the other. If you can only take one number from this post, take that.
Breaking down by model, in the same run:
| Model | prose | long prose | spec |
|---|---|---|---|
| qwen3-coder:30b | 48% | 95% | 95% |
| devstral:24b | 39% | 92% | 85% |
| qwen2.5-coder:32b | 49% | 84% | 81% |
The strongest model ties the two informed arms at 95%. In the other two, long prose is ahead by seven and three points. None of the three put spec ahead. The table aggregates the three repetitions, so it doesn’t prove anything about isolated repetition — stability is shown by the right column of the previous table.
In every model, the big jump continues to be the same: getting out of short prose. The devstral:24b goes from 39% to 92% just because someone wrote what the function does in the corners. Not because someone opened a file with ## Edge Cases in markdown.
Where the average hides what happened
Average is a poor tool for this question. Looking cell by cell in the deterministic run, which are 3 models times 10 tasks, spec and long prose gave identical results in 18 of the 30 cells. In the remaining 12, long prose won 8 and spec won 4.
I went to open one of the identical ones by hand because I didn’t believe it. Task slugify, devstral:24b, temperature 0: both arms produced code byte by byte equal. Two prompts of different size and layout, carrying the same content, collapsing into the same output.
It turns out that, for the model, at that point in the decoding space, table and paragraph were the same request. The information was one. The form evaporated.
formatarBRL: zero in 11 tests, in all three models, by one character
The most interesting finding from the run is qualitative and doesn’t appear in any average. It’s also the best example of what “writing the edge case” means in practice.
In the task formatarBRL, the prose arm scored 0 out of 11 in all three models. Absolute zero, including the happy path. And the cause is the same in all three:
Intl.NumberFormat('pt-BR', { style: 'currency', currency: 'BRL' }).format(1234.56);
// '$ 1,234.56' — visually correct, and fails any string comparison
// the character between "R$" and "1" is not a space:
// code 160 (U+00A0, non-breaking space), not 32
The model wrote the idiomatic solution, the one any human reviewer would approve at a glance. And it fails in assert.equal against 'R$ 1.234,56' with common space because Intl returns an unbreakable space.
When the information is written, and here it fits in one line (“prefix R$ followed by a single common space, do not use non-separable space”), two of the three models start treating the case: either they call Intl and replace the character, or they format by hand. The qwen2.5-coder:32b continues calling Intl without treatment and fails in all three arms, with the spec ahead of it.
This is exactly the “plausible error” I described in July, now at the character level. It’s not absurd code that I would catch in ten seconds of review. It’s the line everyone would write, silently deciding something the request didn’t decide, and that will only appear when the value falls into a CSV or a string comparison on the other side of the system.
If there is one concrete action this post recommends, it’s this: before asking for the function, write the line that formatarBRL needed. One sentence. The character between the R$ and the number. What to do with null. What happens on the last page. You don’t need any table for this. You need the decision to exist somewhere the model will read.
What the spec costs in tokens and time
In the task truncar, the short prose prompt spent 94 input tokens. The same task in spec spent 556, about six times more, for a difference in accuracy that, against long prose, did not appear.
With local models, these 462 additional tokens are prefill time, and it’s little: 4 ms per generation on this card, median of 141 ms vs. 145 ms, five measurements each. With API model they become an account at the end of the month, multiplied by every call in which the spec enters the context.
The cost is the same in both informed arms because the information is the same in both. What the measurement says is that this cost buys the jump from 46% to 90%, and does not buy anything else for being arranged in a table.
In the end, the account that matters is not “spec versus long prose”. It’s “one-sentence request versus I sat down and decided the corners”. The second costs tokens. It also costs your time to think. And it’s the only cost, of those I measured, that returned clear ROI.
Why the bench didn’t go through Aider
In practice, I don’t converse with the model via fetch. I drive the local model through Aider:
# directly in Ollama
aider --model ollama_chat/qwen2.5-coder:32b --read docs/specs/truncar.md src/texto.js
# or through an OpenAI-compatible gateway
export OPENAI_API_BASE=http://localhost:4000/v1
aider --model openai/qwen2.5-coder:32b --read docs/specs/truncar.md src/texto.js
The --read continues to be the detail that matters most. It loads the spec into the context without making it editable. An agent that can edit the spec “resolves” divergence by rewriting the reference, and then I lose the only thing that served as a parameter.
This continues to hold even after this bench. The file didn’t gain generation accuracy for being structured. It gained utility for being an artifact that I control, that doesn’t disappear with the session, and that the agent cannot rewrite to get rid of the problem.
And that’s exactly why this bench didn’t measure through Aider. Aider adds repo map, edit format, and retry. Measuring through it, the number would be from Aider and not from the model. The bench talks to the model via raw API, one turn, no history. It answers “did the model understand the requirement?”, which is not the same question as “did the tool complete the task?”.
Now, where small models really break in daily use in 2025 isn’t writing the function. It’s applying the patch in the search-and-replace format that Aider expects. The block returns with indentation changed or a context snippet that doesn’t match byte by byte with the file, and the edit fails. The code was right all along. This doesn’t appear in any function’s success rate, and it’s most of the real friction for those who use this every day.
Where this measurement doesn’t hold
- Ten small pure functions are not a real repository task. There is no dependency between files, project convention, legacy code to respect. The “out of scope” section, which I called in July the one that pays the most, has almost nothing to do here: there is no neighbor for the model to refactor on the fly. It’s quite possible that the structure pays precisely where this experiment doesn’t look.
- One turn only, without agent loop. No test running, no error coming back, no correction. A good part of the value of an executable acceptance criterion is exactly in the loop that this bench does not have.
- Local models from 24 to 32B. Nothing guarantees that the result transports to a large API model or to a smaller one than these.
- String comparison is fragile, and
formatarBRLproves it. An invisible character zeroed out an entire task. Where my test is too rigid, I fail code that would serve; where it’s loose, I approve code that doesn’t serve. The 3.6 points betweenspecandlong proseare small enough to fit into this fragility. - Three repetitions give an idea of dispersion, not confidence interval. I didn’t do any statistical test, and I’m not going to pretend I did.
Honest reading: what I measured with strength is the effect of writing versus not writing, in pure function, in local model, in one turn. What I didn’t measure is whether the table helps the human not forget the case, whether the file helps in the loop with failing test, or whether in a real repository the out-of-scope section keeps the model. Those questions remain open. I’m not going to push the result where it hasn’t been.
What I changed my mind about, and what to do with it tomorrow
I wrote in July that the spec made the agent more accurate. As I wrote it, it’s wrong. What makes the model more accurate is me deciding and writing what happens when the value is negative, when the input is null, when the page goes past the end. The spec is where I usually do that, not the mechanism that produces the result.
The part that continues to stand from the previous post is the one that matters: the edge cases section is what carries everything. That’s where the 44 points appeared. What falls is the idea that the table, headers, and markdown make a difference to the model. For the model, as far as I measured, they don’t.
So the useful question becomes another. If the structure doesn’t buy generation accuracy, what is the file for? Here I need to be clear because the boundary matters: this is judgment, not measurement. I didn’t test anything that comes below.
- Versioning. Spec glued to the prompt dies with the session. In the repository, it enters the diff and someone reviews the decision before it becomes behavior.
- Review against reference. The question in review stops being “does this seem right?” and becomes “does this match the file next door?”. The second has an answer.
- Reuse between tools. The same file serves Aider, Claude Code, Cursor, and me reading. Prose glued to a chat doesn’t even serve the second time.
- Longevity. The model changes every few months. The file that says what the system does continues to be valid after the change; the prompt tuned for a specific model, not.
The structure serves me and the next person who opens the file. That’s why I continue writing specs. It’s just not what I said it was.
If you want a practical rule coming out of this bench, here it is. Before firing the agent, write the edge cases — in paragraph, list, table, it doesn’t matter to the model. What can’t be done is leave the gap open and hope. After that, if the request is going to live longer than a session, put it in a file with --read and versioning. The first part is what the 46% against 90% paid for. The second is the craft around it, and I still do it, with less illusion about what it buys in pure accuracy.
If you’re going to repeat this, it’s pnpm bancada, with --modelos, --amostras, --temperatura, and --retomar. The tasks are in scripts/bancada-spec/tarefas.mjs, with the three prompts side by side, and that’s the part worth opening before believing anything you just read. Don’t believe technical post numbers. Not even mine. Run.