What we checked
No unit sat on a bench for this review. It was written from NVIDIA's published specifications, set against an RTX 5090 desktop and a Mac Studio M3 Ultra (256 GB) as the capacity-class comparison. The two-box ConnectX-7 cluster, the marquee multi-Spark workflow, is not covered. An earlier version of this section described a stock unit on a bench. There was none.
What you'll feel
The feel below is what the published specs imply. It is not a session on the hardware.
A 70B-class model at Q8 loads and runs. That alone is the headline: 128 GB of coherent LPDDR5X means you stop thinking about quant tier and start thinking about context length. The 5090's 32 GB forces IQ3 or IQ2 on the same model class, every time.
Then you watch the tokens come out and remember the other number. 273 GB/s is the bandwidth ceiling, and decode is bandwidth-bound. Per-token latency on a 70B at long context is slower than a quantized fit on a 5090, and noticeably slower than an M3 Ultra running its 800 GB/s pool. Capacity beats compression; bandwidth beats both.
Multi-model concurrent serving is where the unified pool earns its keep. Two 13B-class models plus an embedding model resident at once, no swap pressure, no pinned-memory choreography. Small-model LoRA fine-tunes complete on overnight runs — tenable, not fast.
The 1 PFLOPS FP4 sparse number is the marketing peak, and nothing here was measured against it.
Setup notes
Out of box on DGX OS, then standard CUDA tooling. The chip is rated at 140 W and ships with a 240 W supply; noise under sustained inference was not measured here. Wi-Fi 7 and 10 GbE on the same chassis is unusual at this size — the 10 GbE saves a thunderbolt-NAS workaround if you're shuttling weights.
Who should buy
- Solo researchers and small teams who need a development box for 70B-class work without leasing cloud hours, and who value desk-sized form factor over absolute throughput.
- People prototyping multi-agent systems where two or three resident models is the workflow.
- Anyone who has to demo native CUDA off a plane.
Who should skip
- Anyone whose workload is one model, one user, throughput-bound. A 5090 tower at the same money decodes faster on what fits.
- A used A6000 build trades MSRP for nearly three times the memory bandwidth — better value for single-model throughput.
- Datacenter buyers stop reading three paragraphs ago.
Bottom line
DGX Spark is a development workstation that happens to fit a 70B model. Buy it for the form factor and the unified pool. Don't buy it expecting H-class bandwidth — the chassis is 1.2 kg for a reason.
