How Much Does It Cost to Program with ChatGPT: $20 Plus vs. API
The choice between subscription and API is wrong in both directions, for the same reason: in chat, each message resends the entire conversation. The cost doesn't grow with the number of questions; it grows with its square.
In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.
en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original
Almost everyone who programs with an online model is paying wrong — and the most interesting part is that the mistake happens in both directions at the same time.
On one side, those who subscribe to US$ 20 and open the chat three times a month. They are paying about seven dollars per conversation and don’t even suspect it.
On the other side, those who migrated to the API because “per token is much cheaper,” turned on the script, sent the entire conversation in each call, and closed the month spending twice as much as the subscription they canceled.
Both mistakes have the same root: nobody does the context math. This post is that math, with prices current as of February 10, 2025 and the arithmetic open for you to redo with your usage.
Today’s Prices
Let’s start with what’s a table fact. Subscriptions first:
| Subscription | Monthly Price (US$) |
|---|---|
| ChatGPT Plus | 20 |
| Claude Pro | 20 |
| Cursor Pro | 20 |
| GitHub Copilot | 10 |
And the APIs, per million tokens:
| Model (API) | Input (US$/1M) | Output (US$/1M) |
|---|---|---|
| GPT-4o | 2.50 | 10.00 |
| GPT-4o mini | 0.15 | 0.60 |
| Claude 3.5 Sonnet | 3.00 | 15.00 |
| xAI grok-beta | 5.00 | 15.00 |
Looking at the second table, the conclusion seems obvious. A 200-token question costs US$ 0.0005 for input in GPT-4o. Half a millimeter of a dollar. With that number in mind, US$ 20 per month becomes forty thousand questions, and the subscription seems like a rip-off.
Add the answer and the number already changes order. A 500-token response costs US$ 0.005 — ten times the question, and the breakdown is clean: the output token costs four times the input, and the response is two and a half times larger than the question. Question and answer together: US$ 0.0105. Now the US$ 20 are worth 1,905 messages, not forty thousand.
It still seems like a lot. This is where the math falls apart.
The Math No One Does
The API has no memory. Each call is independent, and the model knows nothing about what was said in the previous call. If the conversation needs continuity, you carry the continuity: with each message, the entire history is sent again as input, and it’s billed again.
The chat interface hides this very well. You type a line and it seems like you paid for a line. Underneath, the entire conversation went away again.
So the tenth message doesn’t cost like the first one. It costs the first one plus everything that came before. Let’s write this down right.
I call C the initial context — the file you pasted, the system instructions, the stack trace. p is the size of your question and r the size of the model’s response. In turn n:
entrada no turno n = C + (n − 1) × (p + r) + p
entrada acumulada em N = N × (C + p) + (p + r) × N × (N − 1) / 2
saída acumulada em N = N × r
The term that matters is in the middle. It has N × (N − 1), meaning it grows with the square of the number of back-and-forths. Doubling the length of the conversation doesn’t double the cost: it quadruples.
For the math to become a precise number, we need to fix the three parameters. I chose a scenario of a programming session that seems honest to me, and I declare that these are stipulated values, not measured ones: C = 2,000 tokens (a file with about 200 lines pasted at the beginning), p = 200 tokens per question (a paragraph and a small snippet) and r = 500 tokens per response (code with short explanation). Replace them with yours and the structure of the math doesn’t change.
| Back-and-forths | Input in the last turn | Accumulated input | GPT-4o cost (US$) | If it were linear (US$) |
|---|---|---|---|---|
| 1 | 2,200 | 2,200 | 0.0105 | 0.0105 |
| 5 | 5,000 | 18,000 | 0.0700 | 0.0525 |
| 10 | 8,500 | 53,500 | 0.1838 | 0.1050 |
| 20 | 15,500 | 177,000 | 0.5425 | 0.2100 |
| 30 | 22,500 | 370,500 | 1.0763 | 0.3150 |
Thirty back-and-forths — an afternoon session, nothing exceptional — cost US$ 1.08, not the US$ 0.315 that the head math promised. Error of 3.4 times.
Two numbers from this table deserve to be said out loud.
The first one: the thirtieth message alone costs US$ 0.061, against US$ 0.0105 of the first one. The same question, asked at the end of the conversation, costs 5.8 times more than at the beginning. You didn’t change anything in your behavior; you changed the size of what goes along.
The second is worse. Of the 370,500 input tokens billed for the entire session, only 8,000 are text that I typed: the 2,000 from the file plus the thirty questions of 200 tokens each. Everything else is material that had already passed through there before, mine or the model’s. 97.8% of what you pay for input is resent conversation.
This isn’t waste due to provider negligence. It’s how inference works: without history in the prompt, the model has no state. But it’s an account that you can influence, and almost no one tries.
The Break-even Point Has a Formula
With the cost of a session closed, the break-even point is a division:
sessões até empatar com a assinatura = 20 / custo_de_uma_sessão
custo_de_uma_sessão =
( entrada_acumulada × preço_entrada + saída_acumulada × preço_saída ) / 1.000.000
Applying to each model, in the same thirty-turn session:
| Model (API) | 30-turn Session (US$) | Sessions/month until reaching US$ 20 |
|---|---|---|
| GPT-4o | 1.0763 | 18.58 |
| GPT-4o mini | 0.0646 | 309.72 |
| Claude 3.5 Sonnet | 1.3365 | 14.96 |
| xAI grok-beta | 2.0775 | 9.63 |
A little less than nineteen sessions of thirty messages per month, in GPT-4o, and the API has already reached Plus. In Claude 3.5 Sonnet fifteen are enough. In grok-beta, less than ten — and that’s where the US$ 25 credit from beta makes a difference: it covers twelve sessions, which for many people is the entire month without taking the card out of their pocket.
The GPT-4o mini is in another planet: they would be 310 sessions, ten per day, every day. If your work fits into the mini, the subscription discussion simply doesn’t exist.
Translating sessions into hours, which is how most people think about their own usage:
| Duration of a session | GPT-4o (h/month to tie) | Claude 3.5 Sonnet (h/month) |
|---|---|---|
| 30 min | 9.3 | 7.5 |
| 45 min | 13.9 | 11.2 |
| 60 min | 18.6 | 15.0 |
| 90 min | 27.9 | 22.4 |
Less than ten hours per month of dense conversation and the subscription probably doesn’t pay for itself. More than twenty and the API only wins if you do something about the context. It’s exactly this middle range, between ten and twenty hours, that swallows almost everyone — and it’s where the decision is made by detail, not by table price.
What the Subscription Buys and the API Doesn’t Sell
Before treating the choice as an optimization problem, it’s worth listing what the US$ 20 delivers that no token spreadsheet shows.
Spending cap. Twenty dollars are twenty dollars. In the API there is no natural cap: there is a registered card and a limit that you configured, if you remembered to configure it. A poorly closed loop in a night script is a class of error that the subscription simply doesn’t have.
Zero maintenance. No key to store, no .env, no credential rotation, no discovering that the secret leaked in a commit. The cost of this isn’t zero, it’s just invisible.
The interface. Searchable history, file attachment, image upload, voice, resume a conversation from Wednesday. Reimplementing this for personal use is a side project, and side projects of infrastructure are where programmer time goes to die.
In return, Plus has a usage cap that the API doesn’t have — and it moves. The number of messages from GPT-4o per three-hour window varied throughout 2024 and at the beginning of 2025 according to load, according to recurring reports on the OpenAI developers forum. It’s not a number to base an account on: it’s a limit that appears at the wrong time. Those who frequently hit it have the ChatGPT Pro for US$ 200 per month, announced in December 2024 with unlimited access to o1 — a tenfold step, which only makes sense for those who have already proven that the lower step is not enough.
What the API Gives and the Subscription Doesn’t
Three things, in order of increasing importance.
Choosing the model by task. In chat you use the good model for everything, including renaming variables and writing commit messages. In the API, the trivial goes to GPT-4o mini, which costs exactly 6% of GPT-4o — on both sides, input and output. It’s not penny savings: it’s the difference between US$ 1.08 and US$ 0.06 in the same session. The o3-mini, launched at the end of January, falls into the same category for what needs reasoning and not broad knowledge.
Automating. Running on top of a diff, in a commit hook, in batch over one hundred files. Chat requires a person present pasting text; the API doesn’t require anyone. Tools like Aider live off this.
Controlling context. This is the one that pays for the other two, and it’s the subject of the rest of the post.
Cutting Context Is the Lever No One Pulls
In chat, history is sacred: it grows and you watch. In the API, history is an array that you assemble before each call, and you decide what goes in.
Let’s compare three policies in the same thirty-turn session:
| Context Policy | Accumulated Input (tokens) | GPT-4o Cost (US$) | Savings |
|---|---|---|---|
| Full History | 370,500 | 1.0763 | 0% |
| Sliding Window of Last 4 Turns | 143,000 | 0.5075 | 52.8% |
| No History: Only C Plus the Question | 66,000 | 0.3150 | 70.7% |
A sliding window of four turns — which in practice covers almost every programming request, because the current question rarely depends on what was said twenty messages ago — cuts the cost in half and doesn’t change anything in what you type.
And notice the third row: US$ 0.3150. It’s exactly the number from the “if it were linear” column of the first table. The naive math wasn’t wrong by accident — it precisely describes the case where you don’t send any history. All the difference between the naive math and the real math is the history. Who went to the API convinced by the naive math, unknowingly made a promise not to use context. Then they used it.
Prompt Cache: Discount with Toll
There’s a second path that attacks the same problem from another angle: if you’re going to send the same tokens again, the provider can charge less for them. Both big ones had this working by February 2025.
At OpenAI, the prompt cache was announced in October 2024 and is automatic — no code change needed. It’s valid for prefixes with at least 1,024 tokens, and the discount is 50% on the part of the input that has already been seen. The critical point is in the API documentation: the match only happens in exact prefix correspondence, and the prefix in cache usually survives from five to ten minutes of inactivity, up to a maximum of one hour.
At Anthropic, the cache became generally available in December 2024 and is explicit: you mark the cutoff point with cache_control. The price table is different — for Claude 3.5 Sonnet, writing to the cache costs US$ 3.75 per million (25% above normal input) and reading costs US$ 0.30 (one tenth of it). The standard validity period is five minutes, renewed with each read.
Applying this to our thirty back-and-forths. At each turn, what has already been sent before becomes cheap reading and only the 700 new tokens — the previous response plus the current question — enter at full price:
| 30-turn Session | Without Cache (US$) | With Cache (US$) | Savings |
|---|---|---|---|
| GPT-4o | 1.0763 | 0.6413 | 40.4% |
| Claude 3.5 Sonnet | 1.3365 | 0.4138 | 69.0% |
The Claude 3.5 Sonnet goes from more expensive than GPT-4o to cheaper, and it’s not a little: 69% cut. The reason is the asymmetry of the two policies — reading at one tenth of the price compensates with margin the 25% premium on writing, as long as the cache is read several times.
This last condition is the toll, and it has two charges.
The second charge is the clock. A few minutes of inactivity — five in Anthropic’s default, five to ten in OpenAI’s — and the cache dies. Programming isn’t typing non-stop: it’s reading the response, running the test, looking at the log, coming back. If you got up to get coffee between turn eleven and twelve, turn twelve pays full input — and in the case of Anthropic you had already paid 25% more to write a cache that no one read. The best case in the table above is intentionally optimistic; the real case depends on your pace, and mine I haven’t timed yet.
Where Each Choice Wins
Putting it all together, with the thirty-turn session as a unit and three monthly volumes:
| Choice | 5 Sessions (US$) | 20 Sessions (US$) | 40 Sessions (US$) |
|---|---|---|---|
| US$ 20 Subscription (Plus or Claude Pro) | 20.00 | 20.00 | 20.00 |
| GPT-4o mini, Full History | 0.32 | 1.29 | 2.58 |
| Claude 3.5 Sonnet, Cached History | 2.07 | 8.28 | 16.55 |
| GPT-4o, 4-turn Window | 2.54 | 10.15 | 20.30 |
| GPT-4o, Cached History | 3.21 | 12.83 | 25.65 |
| GPT-4o, Full History | 5.38 | 21.53 | 43.05 |
| Claude 3.5 Sonnet, Full History | 6.68 | 26.73 | 53.46 |
| xAI grok-beta, Full History | 10.39 | 41.55 | 83.10 |
The highlighted cell is the one that summarizes the post. Twenty sessions per month — about five hours per week of dense conversation — in GPT-4o with full history: US$ 21.53, more expensive than the subscription this person canceled. Same volume, same model, same work, with a four-turn window: US$ 10.15. The difference between winning and losing isn’t in the provider. It’s in how many times you pay for the same token.
Also notice that the table reverses in the five-session column: there, any API row beats the subscription with margin, and the worst of them — grok-beta with full history — still comes out at a little more than half of US$ 20. Light use, in the API, is almost always the right choice. Heavy use without context control is almost always the wrong one.
The Other Twenty Dollars
It’s worth noting that the choice isn’t binary. GitHub Copilot costs US$ 10 — half of Plus — and Cursor Pro costs the same US$ 20, both with the model inside the editor.
What both sell is precisely the variable that decides the math in this post: they build the context for you. They select snippets from the open file, neighboring files, the repository. You don’t decide what goes in, and you don’t see the token bill.
It’s a legitimate trade-off, and it’s necessary to be honest about its direction: you pay a fixed price not to think about the variable. If the context that the tool chooses is good, you got an advantage. If it’s bad, you have no way of knowing because the bill doesn’t arrive itemized. For predictable use within the editor, it’s the best deal on the list for US$ 10 from Copilot. For work that goes out of the editor — analyzing logs, discussing architecture, batch generation —, it doesn’t replace either chat or API.
What I Still Need to Measure
All the arithmetic above is derived from table price and a session-model that I stipulated. It’s correct, and yet it doesn’t answer the question that matters to me, which is what my actual usage is:
They are five numbers, and I don’t have any of them: how many sessions I do per month, how long each one lasts, how many back-and-forths it has, how many input tokens that gives in a normal work week, and how much actually appeared on the bill.
Without these five numbers, any recommendation I would give would be an opinion with the appearance of math. With them, the formula in the middle of the post answers itself — and that’s why it’s open up there for you to fill in with your usage before I fill it in with mine.
If one phrase from this post remains, let it be this: the price per token is the part of the problem that already comes solved in the provider’s table; what decides your bill is how many times you pay for the same token.
Those who subscribe to US$ 20 bought the right not to think about it — and for heavy and careless use, that right is cheap. Those who went to the API bought control of the variable, but control only becomes a discount when exercised: if the array of messages grows on its own until the end of the conversation, you paid for the key to a safe you never opened.
And the number no one should forget is 97.8%. That’s the fraction of input, in a thirty-message session, that is text the model had already seen. There’s no prompt optimization, model choice, or price negotiation that comes close to moving that fraction. The only thing that exists is deciding what goes in.