Skip to content
Fantástico Mundo de Jon
RSS

Swapping the RTX 5060 Ti for the 5070 Ti doubles bandwidth, not capacity

Both have 16 GB. The same model, quantization, and context fit in both. What doubles is the memory bandwidth, from 448 to 896 GB/s, and it's this that determines tokens per second.

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original

I was on the RTX 5060 Ti 16 GB and the question started to appear frequently: is it worth upgrading to the 5070 Ti? The temptation is obvious. The name is bigger, the launch price is also higher, and the intuition of someone who runs local models is to think that a more expensive card frees up larger models. It turns out that, in this specific pair, this intuition points to the wrong place.

Both cards have 16 GB of GDDR7. So exactly the same model fits, with the same quantization, and the same context. The swap does not change the ceiling of what enters the VRAM. What changes is the memory bandwidth, which exactly doubles: from 448 GB/s to 896 GB/s. And since token generation is limited by bandwidth (each generated token needs to read the model weights in the VRAM once), it is the bandwidth that decides tokens per second.

Whoever swaps expecting to run a larger model is disappointed. Whoever swaps for speed, hits the mark.

What the spec sheet really says

Before talking about models and quantization, it’s worth looking at what NVIDIA published and what can be derived from that without leaving the calculator. There are no benchmarks in this post. No nvidia-smi nor stopwatch. What is below is official specs and arithmetic on top of them.

RTX 5060 Ti 16GB RTX 5070 Ti
Memory 16 GB GDDR7 16 GB GDDR7
Bus 128-bit 256-bit
Memory bandwidth 448 GB/s 896 GB/s
Memory speed 28 Gbps 28 Gbps
CUDA cores 4608 8960
TGP 180 W 300 W
Launch price US$ 429 US$ 749
Launch date 16/04/2025 20/02/2025
Architecture Blackwell Blackwell
Official NVIDIA specs + arithmetic. Bandwidth, CUDA ratio, watts and price are derived from the published numbers; they are not benchmarks.

The exact ratios, just by dividing one column by the other: bandwidth 2.00×, CUDA cores 1.94×, watts 1.67×, price 1.75×. Bandwidth is the only round multiple in the list, and it’s not a coincidence.

Why bandwidth doubles and VRAM doesn’t

Both use memory at the same speed: 28 Gbps. The difference is in the bus. The RTX 5060 Ti 16 GB came with 128-bit. The RTX 5070 Ti came with 256-bit. The memory bandwidth calculation is:

banda = velocidade × largura do barramento / 8

The / 8 transforms gigabits into gigabytes. Substituting:

5060 Ti:  28 Gbps × 128 bit / 8 = 448 GB/s
5070 Ti:  28 Gbps × 256 bit / 8 = 896 GB/s

Same GDDR7, same effective memory clock, double the bus width, double the bandwidth. The capacity remains 16 GB on both sides. There’s no hidden marketing trick here: the RTX 5070 Ti doesn’t load a model that the RTX 5060 Ti 16 GB refused due to lack of space. It only feeds the GPU with the already loaded weights twice as fast.

Token generation is weight reading

During the prefill phase (when the entire prompt enters at once), there’s parallelism left, and the calculation depends more on compute. In generation, the picture inverts. To emit a new token, the model needs to traverse the relevant weights practically entirely, once per token. The batch size is 1. There are no thousands of tokens being processed together to hide memory latency.

If for each token you need to read, say, 8.58 GB of weights, the question becomes: how many times per second can your bandwidth deliver these 8.58 GB? The answer is a division.

teto teórico de tok/s ≈ banda de memória / tamanho dos pesos

This is the theoretical upper limit. It assumes 100% bandwidth utilization, no kernel overhead, no access fragmentation, no sampling cost, and no KV cache getting in line. The actual number is always lower. I didn’t measure it. No real performance numbers appear in this post because I haven’t turned on both cards to compare yet.

Theoretical ceiling at the sizes that matter

I’m using Q4_K_M at 4.9 bits per weight, which is the quantization that actually appears in day-to-day use in llama.cpp and derivatives. The weights below are in decimal GB. The 16 GiB card has 17.18 GB decimals. After accounting for about 1 GB of runtime overhead, the useful budget for weights is around 16 decimal GB. That’s why 27.8B and 32B in Q4_K_M are out of reach on both.

Model Q4_K_M Weights Fits in 16 GiB? RTX 5060 Ti Ceiling RTX 5070 Ti Ceiling
8B 4.90 GB yes 91 tok/s 183 tok/s
14B 8.58 GB yes 52 tok/s 104 tok/s
20B 12.25 GB yes 37 tok/s 73 tok/s
27.8B 17.03 GB no — —
32B 19.60 GB no — —
Official NVIDIA specs + arithmetic. Ceiling = bandwidth / weights, with bandwidth of 448 GB/s and 896 GB/s. It's the theoretical upper limit with 100% bandwidth utilization; it's not a measurement. 16 GiB = 17.18 decimal GB.

Checking the middle table calculations to make sure they’re not floating:

14B em Q4_K_M: 8,58 GB
5060 Ti: 448 / 8,58 ≈ 52,2 tok/s
5070 Ti: 896 / 8,58 ≈ 104,4 tok/s

20B em Q4_K_M: 12,25 GB
5060 Ti: 448 / 12,25 ≈ 36,6 tok/s
5070 Ti: 896 / 12,25 ≈ 73,1 tok/s

In every size that fits on both sides, the RTX 5070 Ti ceiling is double. Not 1.75× (price), not 1.67× (watts), not 1.94× (CUDA cores). 2.00×, because the ratio that rules generation is the bandwidth ratio.

What the swap doesn’t solve

It’s worth repeating in other words because it’s the type of conclusion that slips at the time of purchase.

If your problem is loading a 32B in Q4_K_M, the RTX 5070 Ti doesn’t help. The weights remain at 19.60 GB and the VRAM remains at 16 GiB. If your problem is long context blowing up KV cache on top of a model that’s already stuck at the ceiling, the RTX 5070 Ti also doesn’t help: the cache lives in the same stack of 16 GB. If your problem is offloading to CPU due to lack of space, the swap keeps the offload.

The list of what doesn’t change:

  1. The largest model that fits entirely in the VRAM, at each quantization.
  2. The maximum context before the KV cache pushes layers out.
  3. The need to quantize more aggressively when the model is too large.
  4. The decision between a 20B fully on the GPU and a 32B half on the CPU.

None of this is a function of bandwidth. It’s a function of capacity. Both cards are tied.

What the swap solves

Now the side where the RTX 5070 Ti pays what it costs, at least on paper.

In the same model, same quantization, same context, the theoretical generation ceiling doubles. From ~91 to ~183 tok/s for 8B. From ~52 to ~104 for 14B. From ~37 to ~73 for 20B. In practice, both numbers will be below that, and the measured ratio doesn’t need to come out exactly 2.00×. But the direction is clear: if you’re already satisfied with the model you’re running and only complain about the speed of output, bandwidth is the right lever, and here it doubles.

There’s also the CUDA cores factor, 4608 vs 8960 (1.94×). This weighs more in prefill, in larger batches, and in any load that isn’t strictly token-by-token generation. For the typical interactive use of a local model (chat, agent, completion with batch 1), I continue treating bandwidth as the first term in the equation. Compute is more abundant than memory in this regime. Even so, I won’t pretend that 1.94× cores is irrelevant: on long prompts, prefill should show visible gain, and separate from the generation gain.

What this post compares

Card A
RTX 5060 Ti 16GB
Card B
RTX 5070 Ti
Memory
16 GB GDDR7 on both
Bandwidth
448 vs 896 GB/s
Method
official NVIDIA specs + arithmetic
Benchmark
none — this post is just the math

Price, watt, and the ugly upgrade math

The RTX 5060 Ti 16 GB launched at US$ 429. The RTX 5070 Ti launched at US$ 749. Price ratio: 1.75×. The bandwidth ratio is 2.00×. The CUDA ratio is 1.94×. The TGP ratio is 1.67× (180 W to 300 W). In terms of strictly specification per launch dollar, the additional bandwidth comes at a price slightly below the proportion, and the additional watt comes at a price slightly above. This isn’t a real cost-benefit verdict: street price, availability, and the resale value of the used card factor into the equation, and I haven’t tracked any of those here.

What can be said without leaving the spec sheet:

  • You pay 1.75× the launch price to get 2.00× the bandwidth and 1.94× the CUDA cores, keeping the same 16 GB.
  • You accept 1.67× the TGP. Power supply, case cooler, and noise start to matter more. Power consumption at the outlet and temperature under inference load I didn’t measure.
  • If the RTX 5060 Ti 16 GB is still on your desk, the cost that matters isn’t 749: it’s 749 minus what you recover by selling the RTX 5060 Ti. This residual depends on the market in your country and the day of sale.

Where the theoretical ceiling deceives

I like this bandwidth/weights division because it’s honest about the mechanism. It’s also easy to overinterpret. Three caveats.

First: bandwidth utilization is never 100%. Weight access in inference with batch 1 is regular, which helps, but there are still kernels, layernorm, residual, sampling, and the KV cache itself competing for memory cycles. The real tok/s is clearly below the ceiling on both sides, and doesn’t preserve the 2.00× ratio with two decimal places.

Second: weight size isn’t the only traffic. For each token, you also read and write KV cache. In short context, this term is small compared to a 12 GB model. In long context, it’s not. The ceiling I put in the table ignores this on purpose, to isolate the effect of bandwidth on weights. With large windows, the real generation gain may fall below the ratio of the bandwidths.

Third: lower quantization reduces weights and raises the theoretical ceiling, but doesn’t change the fact that both cards have the same capacity. A 32B in Q3 might eventually fit where Q4_K_M doesn’t. If it fits, it fits on both. The RTX 5070 Ti isn’t a requirement for this hack; it just delivers the token faster after the hack already fits.

The mistake I see most in upgrade discussions

If today you run 14B and want 14B faster, the thesis of the swap holds up on paper. If today you run 20B and want 32B, the thesis falls.

It seems obvious written like this, and yet it’s the most common mistake: mixing the hunger for larger models with the hunger for faster tokens, and thinking that one card only solves both. They are different axes, and this pair of cards separates them with unusual clarity because they tie on one and double on the other.

How I would decide in practice

In my case, the decision starts with a single sentence: am I stuck on capacity or speed?

Stuck on capacity means that the model you want doesn’t load, or loads with offload, or only survives with context too short for use. Then the RTX 5070 Ti isn’t the answer: staying on the RTX 5060 Ti or upgrading to it is the same because the cutoff line is capacity and both have the same. The path is another card with more GB, or another model, or another quantization, or accepting a smaller context.

Stuck on speed means that the right model is already fully in the VRAM and what’s annoying is waiting for the token. Then the 2.00× ratio in bandwidth is exactly the kind of gain I’m looking for, and the RTX 5070 Ti enters the conversation. Before buying, I would still make three measurements on the current card to not swap based on theoretical ceiling:

  1. Tokens per second of generation, in the model and context that I actually use.
  2. Tokens per second of prefill, in the same setup.
  3. VRAM occupied with the full window, not empty.

If the measured generation is already close enough to the theoretical ceiling of the RTX 5060 Ti, the theoretical ceiling of the RTX 5070 Ti becomes a reasonable forecast of gain. If the measured generation is far below the ceiling, the bottleneck may be something else (backend, sampling, CPU in orchestration, PCIe in setup with partial offload), and doubling bandwidth solves nothing while that other bottleneck remains standing.

Both are Blackwell, and that simplifies the comparison

Same generation, same manufacturer, same memory family. This eliminates an entire class of doubt: I’m not comparing old architecture with new, nor GDDR6 with GDDR7, nor software stacks at different stages of maturity. Driver, CUDA, and support in llama.cpp start from the same place on both.

What’s left of difference for local inference is almost all in three axes: bandwidth 2.00×, cores 1.94×, TGP 1.67×. Capacity tied. It’s rare for a card comparison to be this clean.

What I do with this

I don’t swap RTX 5060 Ti 16 GB for RTX 5070 Ti to “start running larger models”. This phrase doesn’t survive the weights table. I swap if the current model already fits, if the generation measured on the RTX 5060 Ti is what’s slowing me down, and if the street price at the time of purchase still makes the bandwidth ratio seem reasonable compared to what I recover from the old card.

In the end, the rule fits in two sentences. Capacity decides what you run. Bandwidth decides the speed. This upgrade messes with the second and leaves the first unchanged, and it’s only with this distinction in mind that the RTX 5070 Ti stops being a talisman and becomes a tool.