The PRD comes first: the spec that prevents the agent from writing the wrong code
The model doesn't err due to stupidity: it errs because it closed the gap left by the prompt on its own, and did so in a plausible way. What changes in the code when the contract comes before the request.
In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.
en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original
I asked the agent for an endpoint with pagination and filtering. A sentence and a half of prompt. It returned in less than a minute: typed, validated, with tests, everything properly indented.
It passed the review because there was nothing visibly wrong. The wrong part were the decisions I didn’t ask for and didn’t read:
pagestarted at0, while the rest of the API starts at1;- sorting was
ORDER BY name, without a tiebreaker — with two suppliers of the same name, pagination repeats one record and skips another; totalcounted the entire table, not the filtered result;sizehad no upper limit:?size=100000pulls the entire table;- page beyond the end returned
404.
None of this is model stupidity. Each item is a gap I left in the request and it filled in alone, with the most common value from what it has seen. The problem isn’t that it makes mistakes — it’s that it makes plausible mistakes. Absurd code I catch in ten seconds; reasonable code that decided the opposite of the rest of the system passes the review and becomes a bug three weeks later.
The bottleneck has moved
In 2023, the bottleneck was the model writing code that compiles. In July 2025, with Sonnet 4 and Opus 4 in Claude Code, with Cursor, with Aider, writing the code for a well-defined task is no longer the hard part. The bottleneck has become the precision of the specification the agent receives.
Rewriting the prompt until it gets it right is a lottery: each round I discover another thing I should have said. Prose has a structural flaw — there’s no place for what I haven’t thought of. A spec has that, and an empty section bothers.
Spec-driven is not PRD-driven
The two documents describe the same delivery, and it’s tempting to conclude that one replaces the other. It doesn’t — they answer different questions, for different readers.
The PRD answers why do it and for whom. The reader is people: those who prioritize, those who sell, those who will inherit this in a year. It exists to align product decisions, and therefore carries business context, success metrics, and discarded alternatives.
The spec answers what exactly the code should do. The reader is the agent. It exists to remove ambiguity from implementation, and within it everything that doesn’t change the code is noise.
| PRD | task spec | |
|---|---|---|
| reader | people who decide | the agent that implements |
| answers | why, for whom, how much it’s worth | what comes in, what goes out, what not to do |
| granularity | a feature | a task, a diff |
| lifespan | as long as the feature exists | updated or deleted along with the commit |
| typical error | too vague to prioritize | too vague to implement |
Combining both into a single file seems like an economy and ends up being expensive. The agent doesn’t separate motivational context from requirements: it weighs everything in the prompt with the same importance. “Suppliers are central to the purchasing flow” is not harmless inside — it’s an input token competing with the line that says total counts the filtered result.
The PRD continues to serve its purpose. It’s just not what I give to the agent. When it exists, the spec is born from it — not as a summary, but as what remains after removing everything that doesn’t become code.
The same request, written in two ways
The task is small and real: list suppliers, with pagination and search. First, how I used to write it.
Cria um endpoint GET de listagem de fornecedores com paginação
e um filtro de busca por nome. Segue o padrão dos outros endpoints.
Each term here delegates a decision silently. “Pagination” is offset or cursor? “Search” is prefix, LIKE or equality? “The pattern of other endpoints” is from which file, if there are three?
Now the same task as a versioned file, in docs/specs/fornecedores-listagem.md.
# spec: GET /api/suppliers
# # contract
GET /api/fornecedores?pagina=1&tamanho=20&busca=&status=ativo&ordem=nome
| parâmetro | tipo | padrão | validação |
|-----------|---------|--------|----------------------------------------|
| pagina | inteiro | 1 | mínimo 1; 0 ou negativo → 400 |
| tamanho | inteiro | 20 | 1 a 100; acima de 100 → 400 |
| busca | string | — | 2 a 80 chars; nome E cnpj; sem acento |
| status | enum | todos | ativo, inativo, todos; fora → 400 |
| ordem | enum | nome | nome, criadoEm, -criadoEm |
The body of the response goes in as a literal example, not as a description:
{
"itens": [
{ "id": "b7c1…", "nome": "Metalúrgica Andrade", "cnpj": "12345678000190",
"status": "ativo", "criadoEm": "2025-07-14T12:00:00Z" }
],
"pagina": 1, "tamanho": 20, "total": 137
}
The unmasked cnpj is in the example — it’s worth more than the paragraph explaining that it comes without a mask.
Here comes the part that decides the result.
# # edge cases
- lista vazia → 200 com `itens: []` e `total: 0`. Nunca 404.
- página além do fim → 200 com `itens: []`. Nunca 404.
- busca com 1 caractere → 400. Não é busca, é varredura de tabela.
- CNPJ mascarado (12.345.678/0001-90) → normaliza para dígitos e compara.
- empate: `ORDER BY nome ASC, id ASC`, sempre. Sem o desempate por id
a paginação repete registro entre páginas.
- `total` conta o resultado do filtro, não a tabela.
- registro com `deletadoEm` preenchido nunca aparece, em nenhum status.
# # out of scope
- não criar índice nem migration; é outra tarefa
- não tocar em GET /api/fornecedores/:id
- não adicionar cache, Redis ou repositório novo — usar o `db`
que já existe em src/infra/db.ts
- não trocar o validador; o projeto usa Zod, mantenha
- sem paginação por cursor nesta entrega
What changes in the code that returns is verifiable line by line. The tiebreaker case, for example:
// without the spec — unstable pagination when there are duplicate names
db('fornecedores').orderBy('nome', 'asc')
.limit(tamanho).offset((pagina - 1) * tamanho);
// with the spec — the tiebreaker was written as an edge case
db('fornecedores').orderBy('nome', 'asc').orderBy('id', 'asc')
.limit(tamanho).offset((pagina - 1) * tamanho);
This is the type of bug that review doesn’t catch. It doesn’t break with ten lines in the table and disappear when you try to reproduce it. It appears on the fifth page, in production.
What doesn’t go into the spec
Out go business justification, persona, adoption metric, and motivational paragraph. They don’t change a single line of the generated code and occupy context that the agent would use better by reading the repository.
The rule: if deleting the sentence doesn’t change the code that returns, it goes out. “Suppliers are central to the purchasing flow” goes out. “total counts the filtered result” stays.
The spec above fits in sixty lines. When I used to write a page and a half, the agent obeyed half — and I didn’t know which half.
Out of scope is the section that pays the most
If I could keep just one section, it would be this one. It’s where the agent invents the most, and invents upwards: creates a repository layer that didn’t exist, adds cache, switches the validator, refactors the neighbor “on the way”, writes migration with a new index.
It’s not sabotage: in the material it was trained on, “good listing code” comes with these things, and without a written boundary it completes the pattern. Pointing to the file that already exists, it uses what’s there — and the diff is small enough for me to really review.
Acceptance criteria need to be executable
“Should work correctly” is not a criterion, it’s a wish. Criterion is the name of the test and the command that runs it.
# # acceptance criteria
testes em testes/fornecedores.listagem.spec.ts, todos verdes:
1. deve_retornar_200_e_lista_vazia_quando_nao_ha_resultado
2. deve_retornar_400_quando_tamanho_acima_de_100
3. deve_paginar_sem_repetir_registro_quando_ha_nomes_iguais
(3 fornecedores de mesmo nome; páginas 1 e 2 com tamanho 2;
os conjuntos de id não podem se intersectar)
4. deve_ignorar_mascara_de_cnpj_na_busca
5. deve_excluir_registro_deletado_logicamente
comando: pnpm vitest run testes/fornecedores.listagem.spec.ts
Each test name came from an edge case in the previous section — the translation is mechanical, and today I write the edge cases already thinking about this.
pnpm vitest run testes/fornecedores.listagem.spec.ts Test Files 1 passed (1) Tests 5 passed (5)
With the criterion like this, the agent closes the loop on its own: runs, reads the failure, fixes, runs again. Without this, I’m the one closing the loop, one round at a time, in the chat.
From endpoint to entire system
An endpoint doesn’t prove any method. The real test was generating a medical scheduling CRUD like this, with a spec per task, never a system spec. A single thirty-page document returns an agent that does everything halfway and nothing to the end.
Scheduling is a cruel domain for prose prompts because almost every rule is a decision no one enunciates:
- two requests for the same time: is the second an error, queue, or fit?
- canceling deletes the record or marks
canceladoEm? (clinical history isn’t deleted — but the agent doesn’t know that) - canceled time goes back to the schedule right away, or only after confirmation?
- does duration come from the procedure or the professional’s schedule?
- what timezone does the time arrive in, and what does the database store?
- holidays and vacation blockades enter the availability calculation?
None of these appear in a request for “make the scheduling CRUD”. All appear in production. And each one that the agent decides on its own it decides well — with the most common answer from training, which in clinics usually is the wrong one: actual DELETE instead of logical cancellation.
The rule that stuck: a task that fits in a reviewable diff, a spec, a test file. When the spec exceeds sixty lines, it’s not a large spec — it’s two tasks.
Where the spec lives
Spec pasted into the prompt dies with the session. The file stays in the repository, versioned with the code, and is cited by path:
# the entire request, after the spec exists
claude "implemente docs/specs/fornecedores-listagem.md"
# in Aider it's the same idea: the spec enters as a read, not as an edit
aider --read docs/specs/fornecedores-listagem.md src/rotas/fornecedores.ts
The --read of Aider matters more than it seems: it loads the spec into the context without putting it in the set of editable files. I’ve seen an agent “resolve” a divergence between spec and code by rewriting the spec — which is strictly the worst possible outcome, because it deletes the only thing that served as a reference.
The CLAUDE.md of the project — .cursorrules, in Cursor — declares the convention once:
Specs de tarefa ficam em docs/specs/. Antes de implementar, leia a spec
citada. Se algo do pedido não estiver na spec, pergunte em vez de decidir.
The last sentence is the one that changes behavior the most: “ask instead of deciding” trades silent decision, which is expensive, for question, which is cheap.
Outside the agent, the spec also pays off: it becomes a PR description with almost no editing, and the review stops being “this seems correct?” to be “does this match the file next door?”. AWS launched Kiro this week, an IDE that makes the requirement file mandatory before generating code — I’m not the only one going this way.
Where this doesn’t pay off
The fixed cost of writing the spec doesn’t disappear when the task is small. I don’t write specs for:
- five-minute task: rename variable, adjust log, bump dependency version. The spec costs more than the task.
- exploration: when I don’t know what I want, the spec is what I’m trying to discover. Then the vague prompt is the right tool — I ask for three approaches and choose one. The spec comes later, from what I chose.
- disposable prototype: if the code dies on Friday, silent decision by the agent costs nothing.
The cut: I write specs when the code will be maintained by someone else — including me in three months — or when the error is silent. Outside of that, prompt and review are enough.
The honest limit
The spec doesn’t make the agent smarter. If I specify wrong, it executes the mistake with precision — it stops being the author of the bug and becomes the executor of mine. It’s better, because the bug ends up in a reviewable file, but it’s not magic.
The spec doesn’t replace the review. It changes what I review: instead of reading the code looking for what’s wrong, I read the diff against the file. It’s an easier question to answer.
And the cost is certain while the gain is still my hypothesis. The spec enters every call as input tokens, and that’s billed — with online model, context becomes a bill at the end of the month. When the model runs on its own board the bill is different, and it’s in the post about VRAM.
I don’t have this number, and I prefer to say that rather than fill out a table. What’s missing to compare, between the two-line prose prompt and the sixty-line versioned spec, are three things: how many rounds until the diff enters without retouching, how many corrections appear after review, and how many input tokens each consumes.
How much time it takes to write the spec above is the number that decides if the method is worth it, and it’s the first one I’m going to time. I haven’t timed it yet.
If I were to start over, I would change the order: I would write the five test names first and let the rest come from them. Top-down, I spend time on the descriptive part — which the agent would infer on its own — and arrive in a hurry at the edge cases, the only section it can’t guess.
The spec doesn’t make the model more capable. It just takes away from it the chance to guess.