Apple Undersold Their Own Chip: M5 Ultra vs M3 Ultra

Comparison|September 21, 2026|By Matthew Moniz|
Share:
comparison11 min read

Full Written Review

Mac Studio M5 Ultra vs M3 Ultra: 15 benchmarks testing Apple's claims. Multi-core hit 1.72x (Apple said 1.3x), the first quad-die Ultra chip, and why the "4.3x AI" number only applies to prompt processing, not text generation.

Apple Undersold Their Own Chip: Mac Studio M5 Ultra vs M3 Ultra

This is the new Mac Studio with the M5 Ultra, and the one in front of me is $12,299 — completely stacked with 256GB of unified memory. They start at $5,500 in the US, but if you want maximum performance with the full core count, you pay for it. At this price, there's really only one question: do Apple's numbers hold up? Apple made three big claims versus the M3 Ultra — 1.3× multi-core, up to 1.8× graphics, and 4.3× AI. So I ran 15 benchmarks across six machines (M1, M2, M3 Ultra, an M5 Max MacBook Pro, and a PC with an RTX 5080) to check.


Quick Verdict

Apple actually undersold the M5 Ultra in several areas — multi-core came in at 1.72× (they claimed 1.3×), and the AI first-token result hit 4.1×. But the "4.3× AI" headline only applies to prompt processing, not text generation (which improved a more modest 1.45×), and that distinction decides whether this machine is worth it for you. If you throw huge prompts, whole codebases, or long-context RAG at local models, it's a game-changer; if you're doing regular video editing and coming from an M3 Ultra, it's not worth the upgrade. Coming from an M1/M2 Ultra, or focused on AI and GPU work, it's a massive leap. New to the Mac Studio? The M5 Max is the smarter buy for most people.


What's Actually Different vs the M3 Ultra

On paper they look close — same GPU core count, same 512GB memory ceiling — but underneath it's a completely different design. Every Ultra chip Apple has ever made was two Max chips fused together with UltraFusion, so macOS sees one big chip. The M5 Max itself is two dies, so the M5 Ultra is four dies — Apple's first quad M-series chip. That's a remarkable feat, because four dies means four pools of memory all pretending to be one. UltraFusion moved about 2.5 TB/s between dies on the M3 Ultra; on the M5 Ultra it's 4.4 TB/s with six times the connection density, and memory bandwidth jumped from 819 GB/s to 1.2 TB/s (up 50%).

The CPU went from 32 cores (24 performance, 8 efficiency) to 36 cores — 12 "super cores" and 24 performance cores. Apple renamed the tiers, and crucially the small cores are no longer "efficiency" cores: there are now 24 real performance cores. It's only 12.5% more cores total, but the little ones got a serious buff. Every GPU core now also has a neural accelerator inside it — a first for an Ultra.

Physically, they're identical (3.7 inches tall, 7.7 × 7.7 inches). The changes are Wi-Fi 6E → Wi-Fi 7, Bluetooth 5.3 → 6, and Gen 6 SSD modules. That's the point: if you've got a rack of these, every mount, dock, and cable already works.


CPU, Graphics, and Creative Benchmarks

Test M3 Ultra M5 Ultra Gain
Cinebench R26 single-core 577 752 1.30×
Cinebench R26 multi-core 10,425 17,943 1.72×
3DMark Steel Nomad 5,532 8,073 1.46×
Blender (3 scenes) 1.67–1.84×
Premiere Pro (PugetBench) 180,370 234,912 1.30×
Photoshop (PugetBench) 14,815 18,360 1.24×
DaVinci Resolve render 5:36 4:54 ~14%

The multi-core result is the fastest I've recorded, and it's 1.72× despite only 12.5% more cores — because Apple didn't just add cores, they replaced 8 weak efficiency cores with 12 real ones. Blender is silly: 8× faster than the M1 Ultra. In DaVinci Resolve, a section that always used to stutter on my M3 Ultra (scrubbing HEIC photos over generators) now scrubs perfectly. A Firefox compile was basically a wash (6 min vs 5), since compiling leans on disk speed and single-thread as much as core count. For gaming (which almost nobody buys a $12,000 computer to do), Cyberpunk 2077 ran 74 → 108 → 166 FPS across the M1/M3/M5 Ultra (Metal FX upscaling, high, frame gen off).


The AI Story: Why "4.3×" Is Only Half True

This is what the machine was heavily advertised for, so I tested it three ways.

  • Time to first token (70B model, 800-token limit, 32,768 context): M3 Ultra 31.6s → M5 Ultra 7.7s = 4.1×, right in line with Apple's 4.3× claim. For reference, the RTX 5080 PC took 30 seconds — slower than a 3-year-old M3 Ultra — not because the GPU is slow, but because the model doesn't fit in 16GB of VRAM and crawls across the PCIe bus into system memory. That's the whole unified-memory argument.
  • Text generation (8B model): M3 Ultra 129 tok/s → M5 Ultra 188 = 1.45×. Not 4.3× — not even close.

These measure two different problems, and understanding it decides whether the machine is worth it. Reading your prompt processes hundreds or thousands of tokens at once — exactly what those new per-core neural accelerators are built for, so you get ~4×. Writing the answer happens one token at a time, and for every token it must read the entire model out of memory, so generation is bound by memory bandwidth, not compute. Bandwidth went up 50%, generation went up 45%. So: if you throw huge prompts, whole documents, entire codebases, or long-context RAG pipelines at local models, going from 31 seconds to 7.7 changes how you work. If you're just chatting with a small model, it's faster — but ask whether that's worth upgrading from an M2 or M3 Ultra.

Image generation (SDXL in Draw Things) went 21.6s → 9.45s (2.29×), and SSD speeds roughly doubled thanks to PCIe Gen 6 (~7,100/8,000 read/write → 14,800/~19,000). Even browser Speedometer jumped (M3 Ultra 40 → M5 Ultra 57), which is really the M5 generation improvement showing up.


Who Is It For?

  • Regular content creation / video editing, upgrading from an M3 Ultra: not worth it.
  • Coming from an M1 or M2 Ultra, or focused on AI / heavy GPU work: a massive upgrade, and the performance is incredible — all while staying quiet, energy-efficient, and drawing far less power than a big desktop PC.
  • New to the Mac Studio world: I'd point most people at the M5 Max — it has the neural accelerators and strong GPU cores, and it's the right SKU for the majority of buyers.

Apple genuinely undersold this chip in raw compute — the surprise is that the headline AI number tells only half the story, and which half matters depends entirely on how you actually use local models.


Where to Buy


Affiliate Disclosure: This article contains affiliate links. If you purchase through these links, we may earn a commission at no extra cost to you.

Published: September 2026

Tags

AppleMac StudioM5 UltraM3 UltraApple SiliconLocal AIBenchmarks2026

Enjoying this video? Subscribe to Matthew Moniz on YouTube for more tech content!

Watch on YouTube