Skip to content

Groq · accelerator

Verified 2w ago

Groq Language Processing Unit

Inference-only ASIC trading HBM for on-die SRAM — deterministic latency for LLM serving.

Stylized line drawing of the Groq Language Processing Unit

Groq's first-generation LPU is a 14nm inference accelerator with 230 MB of on-chip SRAM and ~80 TB/s on-die memory bandwidth. Its PCIe Gen4 GroqCard, no longer sold, delivered up to 750 INT8 TOPS and 188 FP16 TFLOPS, scaling out via 11 RealScale chip-to-chip links.

Specs

compute
Groq LPU (14nm)
form factor
PCIe Gen4 x16

Deterministic-latency inference ASIC. Trades HBM capacity for on-die SRAM bandwidth.

Status, checked 2026-09-24: Groq stopped selling chips to commercial buyers in 2024 and serves its LPUs through GroqCloud. On 2025-12-24 it signed a non-exclusive licensing agreement with NVIDIA for its inference technology; founder Jonathan Ross moved to NVIDIA and Simon Edwards became Groq's CEO. NVIDIA's annual report puts the consideration at $13.0 billion paid at closing plus $4 billion due within a year, and says no customer contracts, products or equity were acquired. NVIDIA's own Groq 3 LPX inference rack, part of its Vera Rubin platform, entered full production on 2026-08-24. In September 2026, DIGITIMES, citing the New York Times, reported that the US Justice Department had opened an antitrust probe into the deal.