The MacBook Pro M4 has more RAM than the 5090 and I stay with CUDA
The 48 GB unified memory in the M4 holds more than the 5090's 32 GB. Nevertheless, vLLM, llama.cpp, and diffusion remain on NVIDIA because decoding is bandwidth and the stack I already use is CUDA.
In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.
en Machine translation by qwen3:32b, reviewed by the author. Read the Portuguese original
Apple just launched the Mac Studio with M5 Max (up to 128 GB) and M5 Ultra (up to 512 GB, 1.2 TB/s bandwidth). The argument is direct: enough unified memory for 70B at home, silent, minimal power draw.
I have a MacBook Pro M4 with 48 GB. That’s more memory than the RTX 5090’s 32 GB in my PC, and more than a 24 GB 4090. I use the Mac daily for work. For LLM and image generation, what’s in the air remains CUDA.
The thing is, fitting and being fast are two separate questions. Apple’s brochure answers the first. The bench answers the second.
What I have on the bench
On one side, the laptop: M4, 48 GB unified memory, Metal, MLX, llama.cpp on Apple backend. On the other, the PC: Ryzen 9 9950X, RTX 5090 32 GB (driver 580), Ollama 0.32 talking to llama.cpp on CUDA, vLLM when the load stops being a chat, ComfyUI with FLUX pinned to the 5090 by UUID.
The 48 GB unified memory is a single well. CPU and GPU see the same RAM. There’s no 32 GB VRAM wall. A 27B or 32B with context that already strains the 5090 enters the Mac without drama. A 70B in Q4 even squeezes in, but it’s bandwidth that’s the bottleneck, not “doesn’t fit”.
This is real. And it’s not what I use to generate tokens.
Fitting and being fast are different axes
I already wrote this in the 5060 Ti / 5070 Ti post: capacity answers “does it fit?”; bandwidth answers “how many tokens per second?”. Batch 1 generation reads almost all weights per token. The theoretical ceiling is bandwidth divided by weight size. Not a measurement. It’s the physical limit.
| Machine | Memory | Bandwidth | What this decides |
|---|---|---|---|
| MacBook Pro M4 48 GB (what I have) | 48 GB unified | ~273 GB/s (M4 Pro) or ~546 GB/s (M4 Max) | Fits more. Slower output. |
| M5 Max 128 GB | 128 GB unified | ~614 GB/s | 70B Q4 comfortably. Bandwidth still ~3× below 5090. |
| M5 Ultra 96 GB | 96 GB unified | 1.2 TB/s | First Apple close to 5090 in bandwidth. Less RAM than Max 128. |
| RTX 4090 | 24 GB VRAM | 1,008 GB/s | 70B Q4 doesn’t fit. Fast where it does. |
| RTX 5090 (in the air here) | 32 GB VRAM | 1,792 GB/s | Token/s ceiling where it fits. CUDA. |
On the 48 GB M4, unified bandwidth is a fraction of the 5090 even in the generous scenario (Max). On the Pro, the fraction is even smaller. 48 vs 32 GB wins the “fits” question. 273 or 546 vs 1,792 GB/s loses the “outputs in time” question.
I won’t put Mac tok/s in this post. I didn’t measure on this machine with the same prompt, same quant, same context as the 5090. Without that, M4 decode numbers would be a guess in disguise.
llama.cpp runs on both sides
Ollama on the Mac works. llama.cpp on Metal too. When I’m away from the PC, this is what I open.
On the PC, the same llama.cpp speaks CUDA. Generation remains bandwidth-limited. Swapping the 5090 for the M4 on this path doesn’t give me quality responses. It gives me extra capacity with a fraction of the bandwidth. You feel it in chat. You feel it more in agents looping tool calls.
The rule I ended up using:
- Fits in VRAM and needs to be fast: 5090, CUDA, llama.cpp/Ollama.
- Doesn’t fit and I accept slower output: Mac, or second GPU, or API.
- Plane, coffee, call: M4, no drama.
The Mac wins 2 and 3. 1 is the path that pays the lunch.
vLLM doesn’t even enter the Mac
vLLM is the server I want when the call stops being a chat and becomes multiple requests at once: PagedAttention, batch, prefix cache. It’s CUDA, with one foot in ROCm. Not Metal. Not MLX.
So “M5 Ultra vs 5090 on vLLM” doesn’t even start. The Mac isn’t in the game. If the product serves LLM with real concurrency, the hardware choice was already made, or you’re paying API.
MLX on Apple Silicon is competent for one user. It’s not the same software. “OpenAI-compatible” on paper doesn’t erase scheduler, prefix cache and tensor parallel. I won’t rewrite the stack in MLX just because Apple increased bandwidth.
Image isn’t memory either
Image generation here is FLUX on the 5090, ComfyUI isolated by UUID on the GPU, on the same host as Ollama. I’ve had thumbnails coming out via cloud with the 5090 idle. That was a file in the wrong place, not lack of Mac.
ComfyUI, ControlNet, the weights the community actually uses: CUDA. MPS on the Mac exists. Existing isn’t the pipeline I already operate and measure. 48 GB unified fits a FLUX that 24 GB of 4090 also fits. The 32 GB 5090 fits with more room for LoRA. None of these three cases make me prefer Metal. What makes me prefer NVIDIA is the minute between prompt and image in the air, and the fact the rest of the house is already on this track.
What the M5s change in the math
M5 Max 128 GB: the first notebook where 70B Q4 with long context is comfortable, not “fits if you pray”. The generational jump Apple is selling is prefill (pasting an entire repository), via Neural Accelerators on each GPU core. Decode rises by 20-30% over M4 Max, following bandwidth from ~546 to ~614 GB/s. Still far from 5090 on models that already fit there.
M5 Ultra 96 GB: 1.2 TB/s. Then the physics of decode gets close. 96 GB, irony, is less memory than the Max 128 GB. You pay for Ultra on bandwidth, not tank. 256 and 512 GB are another price bracket.
None of this gives me vLLM. None of this gives me the ComfyUI already in the air. None of this changes the fact that, on models that already fit in the 5090, CUDA wins ugly.
What this post compares
- Laptop
- MacBook Pro M4, 48 GB unified
- PC
- Ryzen 9 9950X + RTX 5090 32 GB
- LLM on PC
- Ollama 0.32 / llama.cpp CUDA, vLLM
- Image
- FLUX + ComfyUI on 5090
- Bandwidth
- Apple and NVIDIA specs, not M4 tok/s
- New Macs
- M5 Max and M5 Ultra, announced 25/08/2026
Where the Mac truly wins
Work. Screen, battery, silence, IDE, call. 48 GB is enough for this. Not this debate.
Models that don’t fit in 32 GB. 70B in Q4, absurd context on top of a 27B, two models at once. Then unified memory stops being brochure and becomes a ceiling the 5090 doesn’t have.
And the lap. I won’t carry the 5090 on the plane.
The mistake is treating the Mac as a GPU with huge RAM. It’s not. It’s an SoC with a single memory well, smaller bandwidth (except Ultra), and another ecosystem. This is great for fitting. It’s bad for generating.
I like the M4. I’ll keep walking with it. The M5s are Apple’s best argument yet for on-device AI. I’ll stay on the 5090, and a 4090 on the same CUDA track, for token generation, batch, and image.
Unified memory solves “doesn’t fit”. CUDA solves “doesn’t have time”. In my flow, the second sentence appears more often.