Skip to content

Brief · 5 October 2026

What changed

Qwen 3.8 Flash Next (125 B) was benchmarked on a consumer RTX 4090 and hit a peak of 100 T/s token throughput, a record for a model of this size on desktop hardware.

One number

100T/s

Peak token throughput for a 125 B model on an RTX 4090

source ↗

Still vapor

NVIDIA’s Display Config Server is marketed as a latency‑killer for AI‑driven graphics, yet the Wayland vs. X.Org latency tests published by Phoronix show only a few‑millisecond edge, far short of the sub‑millisecond gains implied for high‑performance AI pipelines.

The most concrete shift today comes from a community‑run benchmark that pushed Qwen 3.8 Flash Next, a 125‑billion‑parameter model, on a single RTX 4090 and measured 100 T/s token throughput. The result proves that consumer‑grade GPUs can now sustain data‑center‑level token rates for large language models, reshaping the make‑vs‑buy calculus for labs that rely on on‑prem inference rigs. Practitioners note that the test used the latest TensorRT‑LLM optimizations and a 48 GB HBM2e memory configuration, but the raw throughput figure is what matters for capacity planning.

At the same time, NVIDIA’s recent Display Config Server announcement at XDC 2026 was accompanied by a Wayland vs. X.Org latency study from Phoronix. The paper reports only a modest 2‑3 ms improvement, contradicting the vendor’s hype of “near‑zero latency” for AI‑augmented display pipelines. For operators, the takeaway is that the claimed latency edge does not yet translate into measurable gains for AI workloads that are already latency‑sensitive.

AMD’s latest DRM‑Next pull for Linux 7.4 adds a handful of bug fixes and a new GPU enablement, but without performance numbers it remains a maintenance update rather than a capacity‑shifting event. In short, today’s headline is the RTX 4090’s 100 T/s token record, while NVIDIA’s latency promises fall short of the data.

Composed by the MadCoolStuff editor pipeline · Groq · openai/gpt-oss-120b · 2026-10-05

Tags

What we read