Skip to content

Apple · workstation

Verdict · buy-if

Mac Studio M3 Ultra: the patient context machine

Apple's M3 Ultra Mac Studio fits 100B+ class models a 5090 has to page or skip. First-token latency lags; total throughput wins when the model doesn't fit elsewhere. A specialist box, not a generalist.

Product
Apple Mac Studio (M3 Ultra)
Published
2026-05-01
Price
$9,499
Score
8 / 10
8/10
Stylized line drawing of the Apple Mac Studio (M3 Ultra)

Pros

  • Half a terabyte of unified memory in a desktop that draws less than a gaming GPU
  • Runs 100B+ MoE models at usable quants without offload juggling

Cons

  • Memory bandwidth tops out at 819 GB/s — well under a 5090's, and you feel it on first token
  • Image and video gen workflows are second-class on Apple Silicon; expect minutes where CUDA does seconds

Verified numbers

verified 2026-09-24

  • unified memory max (GB)

    512

  • bandwidth (GB/s)

    819

  • msrp (USD)

    9,499

  • comparator 5090

    5,090

  • comparator 5090 vram (GB)

    32

  • comparator h100 label

    H100

Superseded. Apple replaced this Mac Studio with M5 Max and M5 Ultra models on 2026-08-25, and the M3 Ultra build is no longer sold new. The review below is about the M3 Ultra as it launched, at the $9,499 the 512 GB configuration cost.

What we checked

No unit sat on a bench for this review. It was written from Apple's published specifications, and it weighs the workloads where the M3 Ultra has a structural argument — capacity per dollar, capacity per watt — and skips the ones where the answer is already known to be no. An earlier version of this section was headed "What we tested". Nothing was.

  • 70B-class dense at Q8 — full context window, single-user chat, no quantization-quality compromise.
  • 100B+ MoE at Q4 (Qwen3-235B-class) — the workload a 32 GB 5090 cannot run without paging or aggressive offload.
  • Mixed local agent loop — long-context retrieval over a working set that wouldn't fit in 32 GB without sharding.

What you'll feel

The feel below is what the published specs imply. It is not a session on the hardware.

First token is slow. The bandwidth ceiling is 819 GB/s, and at 100B+ active parameters that ceiling is the bottleneck — you'll wait noticeably longer for the first token than on a 5090 running anything that fits in 32 GB. If your workflow is interactive — code completion, chat-style iteration on small models — buy the 5090 and stop reading.

Where the M3 Ultra earns its keep is the moment your model doesn't fit on the GPU. A 5090 forced to spill weights to system RAM collapses to single-digit tokens/sec. The Mac doesn't spill, because there is no spill — the model is already in unified memory. Steady-state throughput on a 100B-class MoE is competitive with a 5090 running the same model under offload, and the Mac gets there without thermal events, fan noise, or PCIe juggling. It's a different shape of fast.

Setup notes

  • llama.cpp Metal backend is the reliable path. MLX is faster on supported model families but less universal.
  • The 512 GB SKU is the only configuration that justifies the contrarian thesis. 256 GB and below, the value math gets harder.
  • Sequoia's memory pressure UI lies a little — unified memory accounting includes weights resident for inference. Watch vm_stat, not Activity Monitor.

Who should buy

  • Researchers running 100B+ class models locally for privacy or iteration speed reasons, who can tolerate slow first-token in exchange for not renting H100 hours.
  • Engineers whose working set is the model size, not the latency floor — long-context RAG, batch evaluation, dataset distillation.

Who should skip

  • Anyone whose primary workload is 70B and below and prioritizes time-to-first-token. A 5090 is faster and a third the price.
  • Image and video generation users. ComfyUI, SDXL training, video diffusion — Apple Silicon is workable but consistently behind CUDA.

Bottom line

The M3 Ultra Mac Studio is not the obvious choice and does not pretend to be. It is the rig you specify when the model is the constraint and the latency budget is generous. For everyone else, a 5090 wins on price, on first-token speed, and on ecosystem maturity. Pick the tool that matches the workload — and if your workload is a 100B+ model you need to run today on hardware you own, the answer is on a desk in Cupertino.