Skip to content
Fantástico Mundo de Jon
RSS

A complete API through chat: what GPT-4o gets right and where it misleads you

In November 2024, no agent edits my files: the API exits the chat in blocks that I paste by hand. GPT-4o gets the structure right and misleads me in four places — the worst is the test that passes without testing anything.

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original

The chat writes your entire API. It just never knows that it exists.

This is not a metaphor. In November 2024, there isn’t something that opens my repository, reads the files, decides what to change, applies the changes, runs the test, and comes back with the result. There is Cursor, which edits the file that is open when I accept the suggestion. There is Aider, which applies the patch and commits what I told it to apply. There is Copilot completing the line. But to build something from scratch, the flow that I — and almost everyone I know — actually use is as dumb as possible: I describe, the model returns a block, I paste it into the editor, run it, and come back with the error pasted again.

Copying and pasting is not a detail of the method. It is the method.

The loop is always the same. I describe, it returns, I paste, TypeScript complains, I copy the error back, it fixes that line and reintroduces something two blocks above. I am the one who closes the loop, one round at a time, with the clipboard. And the detail that explains half of the problems in this post: the model never sees the result of running what it wrote — it sees my transcription of the result, filtered by what I thought was relevant to paste.

I did this from start to finish on a REST API for orders. Short verdict: GPT-4o gets what has a predictable shape right and messes me up in four places. In all four, the error only appears after pasting.

What was built · November 2024

Task
REST API for orders, 4 endpoints
Stack
Node 20 · TypeScript 5.6 · Fastify · Knex · Postgres
Tool
ChatGPT Plus, GPT-4o, browser window
Method
copy and paste; the model does not read the repository
Fixed preamble
159 tokens per message

None of these tools opens the repository on its own

It’s worth situating what exists and how much it costs, because in a year this list will seem prehistoric and the price is the only part that doesn’t depend on my opinion.

Tool Price in Nov/2024 What it does with my file
ChatGPT Plus $20/month nothing — I paste and copy back
GitHub Copilot $10/month completes the line; in chat, suggests and I apply
Cursor Pro $20/month edits the open file when I accept
Aider open source applies the patch and commits, one request at a time
Prices from the current table as of November 2024, individual plan. Aider is open source: what you pay for is the model to which you point it. None of the four decides on its own what to change in the repository — the decision and application remain mine.

Cursor and Aider shorten copy-and-paste, each in their own way, and are worth what they cost. But the topic of this post is the mode that is still the most common for starting a project from scratch: a browser tab, a prompt, and the clipboard.

Asking for the contract before the code changes what comes back

The first request I made was the natural one, the one anyone types without thinking.

Faz uma API REST de pedidos em Node com TypeScript.
Precisa criar, listar, buscar por id e cancelar.

It returned a single file, nice and immediately useless. The part that matters is this:

// vague request response: route, business rule and SQL in the same place
app.post('/orders', async (req, res) => {
  const { customerId, items } = req.body;
  const total = items.reduce((s, i) => s + i.price * i.qty, 0);
  const { rows } = await pool.query(
    'INSERT INTO orders (customer_id, total, status) VALUES ($1,$2,$3) RETURNING *',
    [customerId, total, 'pending'],
  );
  res.status(201).json(rows[0]);
});

This runs. And it has four problems, only one of which is style-related. The domain came in English, while the rest of my system speaks Portuguese — a cosmetic detail that turns into chaos when both coexist. There is no validation at all: items can be undefined and the reduce blows up with 500. The response returns the raw line from the database, with internal columns included. And the fourth one is the real deal: the total is calculated based on the price that the client sent in the request body. Whoever calls the API decides how much they pay. The model didn’t invent this out of malice; it completed the most common pattern of “cart coming from the front,” and I didn’t say where the price comes from.

Then I changed the first message. Instead of asking for code, I asked for the contract.

Antes de qualquer código: escreva o OpenAPI 3.1 de /pedidos.
Só o contrato — caminhos, parâmetros, corpos, códigos de erro.
Nenhuma implementação. Vou revisar e devolver corrigido.

What comes back is a YAML that I read in two minutes, and it’s in those two minutes that the important decision happens.

paths:
  /pedidos:
    post:
      summary: Cria um pedido
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required: [clienteId, itens]
              properties:
                clienteId: { type: string, format: uuid }
                itens:
                  type: array
                  minItems: 1
                  items:
                    type: object
                    required: [produtoId, quantidade]
                    properties:
                      produtoId: { type: string, format: uuid }
                      quantidade: { type: integer, minimum: 1 }
      responses:
        '201': { description: Criado }
        '409': { description: Produto sem estoque }
        '422': { description: Corpo inválido }

Notice what is not in the request body: preco. That line was not in the first version of the YAML — I deleted it, and returned saying that price comes from the products table, always. It was a ten-second edit in a text file, and it kills the billing bug before there is a single line of code. Reviewing prose is expensive and reviewing the contract is cheap: the contract has a defined place for each decision, and an empty place bothers.

The error codes at the end of the YAML are also mine, not his. The first version responded 400 to everything; 409 for product out of stock and 422 for invalid body were two lines that I changed before there was any implementation. Discussing this in the contract costs ten seconds. Discussing it later, with a handler already pasted into the editor and a green test around it, costs an entire conversation — and usually I let it pass because it’s already working.

The second gain is more prosaic and perhaps bigger. The YAML becomes the artifact that I paste at the top of every following message. In a conversation that has no access to my disk, this file is the only long-term memory that exists.

What it gets right is everything that has shape

Being fair before being annoying: the structural part GPT-4o delivers well, quickly, and almost always on the first try.

I asked for the validation scheme based on the contract — literally “generate the Zod that validates the body of POST /pedidos from the OpenAPI above”.

// schemas/request.ts
export const criarPedidoSchema = z.object({
  clienteId: z.string().uuid(),
  itens: z
    .array(
      z.object({
        produtoId: z.string().uuid(),
        quantidade: z.number().int().positive(),
      }),
    )
    .min(1),
});

export type CriarPedidoDTO = z.infer<typeof criarPedidoSchema>;

It’s correct. The only thing I changed was the error message, which came in English. Separating route service and repository, writing the DTO, deriving the type from the schema, assembling the try/catch that translates validation error into 422, writing the test file with the describes in place: this entire layer comes out ready and compiles. It’s boring and repetitive work, the kind that consumes an entire afternoon of typing without requiring any decision. It comes out in seconds.

This is real and it’s not little. It’s also not the hard part. The hard part is everything that depends on something that is on my disk and not in the conversation — and that’s exactly where the second half of the post begins.

The library it writes is the one it saw, not the one in my package.json

The model writes the API of the version that appeared in its training. Not the version installed here. It has no way of knowing which one it is because it never opened the file.

I asked for the connection to Postgres, with pool and timeout, and pasted what came back.

// request: "Postgres connection, with pool and connection timeout"
const pool = new Pool({
  connectionString: process.env.DATABASE_URL,
  max: 10,
  idleTimeout: 30000,      // nesta versão a opção chama idleTimeoutMillis
  connectionTimeout: 2000, // e aqui, connectionTimeoutMillis
});

Neither of those two lines breaks anything. Configuration object accepts unknown key without complaining: the pool goes up with defaults, and the two timeouts I asked for simply do not exist. In development this is invisible because the database responds in 3 ms and nothing ever waits. It appears on the day when the database becomes slow in production and the request hangs until someone notices.

This is the expensive case. The cheap case — import that no longer exists, renamed function, changed signature — TypeScript catches in two seconds and you fix it without thinking. What hurts is the silent variant: code that runs, looks like what you asked for, and does something else.

The mitigation is silly and works well: paste the versions along with the request, every time it touches a library.

pnpm list --depth 0

dependencies: fastify 4.28.1 knex 3.1.0 pg 8.13.1 zod 3.23.8 jsonwebtoken 9.0.2

the block I paste in the prompt whenever the request touches a dependency

With this pasted, the hit rate visibly increases. And even so it keeps making mistakes when the thread stretches because that block gets left behind in the conversation while the current request is at the front. It’s not forgetfulness in the human sense; it’s the relative weight of what is near the end of the prompt.

The migration doesn’t match the entity because they were two conversations

The type of the entity came early, along with the service, and is exactly how I wanted it.

// types/request.ts — generated in message 6
export type Pedido = {
  id: string;
  clienteId: string;
  total: number;
  status: 'aberto' | 'pago' | 'cancelado';
  criadoEm: Date;
};

Eight messages later, with the service already running against a mock, I asked for the Knex migration to create the table.

// migrations/20241112_create_requests.ts — generated in message 14
export async function up(knex: Knex) {
  await knex.schema.createTable('pedidos', (t) => {
    t.increments('id');
    t.integer('cliente_id').notNullable();
    t.float('total');
    t.string('status').defaultTo('pending');
    t.timestamps(true, true);
  });
}

There are five divergences in six lines. The id became sequential integer, and the type says string because it’s uuid. The clienteId became cliente_id without anyone configuring case conversion, and without a foreign key. The total is float, which is the classic way to lose a cent in rounding — money is an integer in cents, and that I had said in message 6. The status gained default 'pending', in English, outside of the union of three values in Portuguese and without any restriction in the database. And timestamps() creates created_at and updated_at, while the type expects criadoEm.

None of this breaks the build. TypeScript doesn’t talk to Postgres: the entity type is a promise that no compiler checks. You find out on the first real query, and you find out in pieces — first the id, an hour later the status, and the float only when someone sums ten orders and the total closes with a crooked cent.

The cause is always the same. Between message 6 and 14 came dozens of code blocks, and the model does not have the file tipos/pedido.ts open next to it. It has the conversation history, competing for attention with everything else. The correction is to paste the type along with the migration request, and the lesson is that “do you remember the type you yourself wrote?” is a question that doesn’t make sense to ask.

Better than correcting is reversing the order: asking for both in the same message, one below the other, so that it writes the table looking at the entity. That’s what I started doing, and the problem practically disappeared. It’s still my job to decide that these two things need to be born together — there’s nothing in the conversation that informs the model that there is a Postgres on the other side waiting to match that type.

The auth comes right from the outside and wrong from the inside

I asked for an authentication middleware with JWT. This came back, and this is where the method becomes really dangerous.

// request: "JWT authentication middleware"
export async function autenticar(req, reply) {
  const token = req.headers.authorization?.split(' ')[1];
  if (!token) return reply.code(401).send({ erro: 'sem token' });

  const payload = jwt.decode(token);
  if (!payload) return reply.code(401).send({ erro: 'token inválido' });

  req.usuario = payload;
}

Before the real problem, a less severe warning: the first version came with Express signature, (req, res, next), with res.status().json() and next() at the end. The project is Fastify, and that’s written on the dashboard up there, but I didn’t repeat it in the prompt of that message. The model completed with the most common format on the internet. This error is cheap: it doesn’t run, you find out in thirty seconds. Above is the version already converted to Fastify preHandler — which is where the expensive error lives.

All observable behavior is correct. Without header, 401. Malformed header, 401. Random string, 401. Valid token, passes and the req.usuario arrives filled in the handler. You test it in Postman, it works in all four cases, and you move on.

jwt.decode does not verify signature. It does base64 and returns the content. Anyone who builds a JSON with another user’s id, encodes it, and sends it, enters as that user. And exp is also not checked, so token expired in January continues to be valid in November. The correction is one word: jwt.verify(token, secret, { algorithms: ['HS256'] }). The algorithms option at the end is not decoration — without it, the verifier accepts the algorithm that the token itself declares, which is an entire class of free attack.

One word and one option. In code review, this file passes: it has the exact format of an auth middleware, with the right names and 401s in the right places.

The test that passes and doesn’t test anything

This is the central problem, and it’s what made me write the post.

After the order service was up, I asked for the most natural thing in the world: “write the tests for this service with Vitest”. A green file came back.

// tests/request.service.spec.ts — what came back
import { describe, it, expect, vi } from 'vitest';
import { criarPedido } from '../src/servicos/pedidos';

vi.mock('../src/servicos/pedidos', () => ({
  criarPedido: vi.fn().mockResolvedValue({ id: '1', total: 100 }),
}));

describe('criarPedido', () => {
  it('deve criar um pedido', async () => {
    const pedido = await criarPedido({ clienteId: 'c1', itens: [] });
    expect(criarPedido).toHaveBeenCalled();
    expect(pedido.total).toBe(100);
  });
});

Read the third line after the imports. The test mocks the module that is under test. The criarPedido that runs inside the it is not my function: it’s a vi.fn() that returns total: 100 because the line just above told it to return. I could delete the entire src/servicos/pedidos.ts and this file would still be green. The assertion toHaveBeenCalled asserts that the function that the test itself just called was called. And the itens: [] in the call is exactly the case that the contract says to respond with 422 — the test doesn’t notice because nothing real executed.

pnpm vitest run

Test Files 1 passed (1) Tests 1 passed (1)

green, without touching a line of my code

The second variant is more subtle and appeared many more times. Here the mock is in the right place — the repository, which is the dependency — and the problem has moved.

// the mock is correct; the assertion is that it doesn't look at anything
const repo = { salvar: vi.fn().mockResolvedValue({ id: 'p1' }) };

it('calcula o total do pedido', async () => {
  await criarPedido(repo, {
    clienteId: 'c1',
    itens: [{ produtoId: 'x', quantidade: 3 }],
  });
  expect(repo.salvar).toHaveBeenCalled();
});

This runs my real code, which is already a step up. But the test name promises the calculation of the total and the assertion doesn’t look at the total. If criarPedido multiplies by quantity, ignores quantity, adds instead of multiplying, or returns NaN, the test stays green in all four cases. toHaveBeenCalled without arguments is the weakest assertion in the suite: it asserts that execution passed through there, which a console.log would also assert.

The fix is one line, and the difference between the two lines is the entire post:

expect(repo.salvar).toHaveBeenCalledWith(
  expect.objectContaining({ total: 2970 }),  // 3 × R$ 9,90 em centavos
);

Why does it do this? It’s not laziness, it’s the target. Green suite is the only signal of “I’m done” that exists in that conversation, and there are two paths to get there: make the code work or make the test not look. The second is shorter and has exactly the same appearance. Added to this, the model saw a huge amount of bad tests in training — toHaveBeenCalled is probably the most written assertion in JavaScript history.

And it doesn’t help to look at coverage, which is the natural reflex at this point. The first file, the one that mocks the module under test, marks zero line covered by the service and denounces itself. The second variant does the opposite: runs my entire code, covers each line of it, and only doesn’t check the result. High coverage with an assertion that accepts any number. The report measures whether the line executed, never whether someone looked at what it produced.

The ritual I started doing, always, right after any test comes back: ask it to break the function on purpose and say which test fails. If the answer is “none”, the file is decoration. Even better is to do this manually — change a * to + in the service and run. It takes fifteen seconds and is the only way to know if you have a test or a green ornament.

Each response brings back what I had already fixed

The structural cost of copy-and-paste is not the time of Ctrl+C. It’s that the model doesn’t see the current state of the repository, so every response is generated from a world that stopped at the last thing I pasted — and in it my manual corrections never happened.

The typical case: I fix the jwt.decode manually in the editor. Three messages later I ask for a change in the middleware error response. It returns the entire file, politely rewritten, with the jwt.decode back in place. This is not model regression; it’s that my correction never passed through it.

Against this, I put together a fixed preamble that goes at the top of every message asking for code.

ESTADO (não altere nada fora do que eu pedir)
- Node 20, TypeScript 5.6, ESM, Fastify, Knex, Zod, Vitest
- Domínio em português: pedidos, itens, clientes. Nunca order/item.
- id é uuid gerado no banco. Nunca increments.
- Dinheiro é inteiro em centavos. Nunca float.
- Erro sai em ProblemDetails (RFC 7807).
- Contrato vigente: o OpenAPI da mensagem 1 (/pedidos e /pedidos/{id}).
- Já corrigido, não regrida: jwt.verify com algorithms HS256.

PEDIDO
<uma tarefa, uma só>

ARQUIVO ATUAL
<cola do arquivo que vai mudar, inteiro>

It costs 159 tokens per message — I counted with the GPT-4o tokenizer over the block above — and is the cheapest thing in the process. At $2.50 per million input tokens, it’s $0.0004 per message: thirty messages of preamble cost a little more than one cent. The line that pays the most is the last one in the state block: “already corrected, do not regress” is what made the same three bugs stop coming back with each rewrite. And it doesn’t solve everything: past a certain point in the same thread, it starts inventing again. When this begins, I open a new thread and paste contract plus state; it’s cheaper than fighting history.

It’s worth noting what this ritual is, looking from the outside: me typing manually, every message, a worse and manual version of what the tool should read from disk itself. It’s an index maintained by a human with the clipboard.

ChatGPT, Claude, and Grok on the same task: what can be asserted

I did parts of the same work in all three. Before impressions, prices — which are facts, unlike the rest of this section.

Model (API) Input per 1M Output per 1M
GPT-4o $2.50 $10
GPT-4o mini $0.15 $0.60
Claude 3.5 Sonnet $3 $15
xAI grok-beta $5 $15
Prices from the current table as of November 2024, per 1 million tokens. The xAI API entered public beta on November 4, 2024, two weeks ago, with $25 per month of credit — that's what it is, and not technical curiosity, that put grok-beta in many people's comparison, including mine. This work was done on ChatGPT Plus at $20/month, not through the API; the table is here because it's the cost to automate the same flow.

GPT-4o. Fast and economical in response. Tends to return the snippet and write “the rest remains the same”, which is great when I know where to fit and terrible when the file is new. It was the one that most consistently transformed my prose into usable OpenAPI on the first try, for this task.

Claude 3.5 Sonnet. Returns a more complete file and writes more. Held up better to long pastes — entire file plus contract plus error — without losing what was at the beginning. It’s also the one that returned the most entire file instead of the snippet, which is exactly the behavior against which the previous section warns. Maintained the domain in Portuguese after I asked once.

grok-beta. It worked. I didn’t observe anything that would make me switch tools for this task, and the reason it’s here is the beta credit, not a hypothesis of mine about the model.

Two footnotes that only make sense today. The GPT-4o mini, at $0.15 and $0.60 per 1M, is cheap enough to treat boilerplate — DTO, mapping, test skeleton — as disposable, on the day I switch the browser for a script. And o1-preview is available, but it’s not for this: it thinks before responding, and what this task asks is the opposite, dumping file after file while I paste. I reserved it for the isolated difficult question, not for the loop.

What can be asserted with more security is boring: the distance between the three is much smaller than the distance between “I wrote the contract first” and “I didn’t write it”. Switching models changed the style of what comes back. Writing the contract changed what comes back.

What separates the two halves

Three out of four places where the chat messed me up have the same shape: the library version, the migration, and the green test. All depend on information that is on my disk and not in my prompt: the package.json, the type file written eight messages before, the real behavior of the function under test. Where the model can generate from what I wrote, it’s good and it’s fast. Where it would need to read something I didn’t paste, it completes with the most likely, and the most likely is always plausible, which is much worse than absurd. Absurd code I catch in ten seconds.

Auth is the fourth place and doesn’t fit this pattern. There wasn’t missing information from me: there’s no file on my disk that answers whether jwt.decode verifies signature. The model lacked choosing the safe default instead of the most frequent one. That’s why it’s the most dangerous of the four — pasting more context doesn’t solve it, and it’s exactly the knowledge you outsourced by asking for the middleware ready-made.

So the work today is deciding what goes into the text. Asking for the contract before the code is that and nothing else: moving into the prompt the decision that, if left out, the model takes alone and in silence. Pasting versions is that. The state block at the top of the message is that.

Part of this ritual will become obsolete as soon as the tool opens the repository on its own, and that’s good. The other part won’t: a test that mocks the thing under test remains green regardless of who opened the file, and no disk access fixes an assertion that doesn’t look at the result.

Meanwhile, the role is clear. In this conversation I act as compiler and linter, and that’s fine. What really bothers me is acting as memory.