Skip to content
MCS · Models · 42

The Lattice.

The latest models, plotted in benchmark space.

Rate the field — ranked, tiered, by recipe →

Pick any benchmark for each of three axes. Drag to orbit. 24 cloud + 18 open-weight models, every number traced to its source.

Snapshot · verified 1w ago · 2026-08-17

X axis↑ higher better
Y axis↑ higher better · tok/s
Z axis↓ lower better · $/M
Color by

The 3D lattice needs a desktop browser with WebGL. The full matrix is in the table below.

34 of 42 plotteddrag to orbit · scroll to zoom · click a node
Hover a node to read it; click to pin. Pick any benchmark for each of the three axes — the cube re-plots live.
Show
Alibaba (Qwen)AnthropicDeepSeekGoogle DeepMindMiniMaxMoonshot AIOpenAIxAIXiaomiZ.AI (Zhipu)

Reading the axes

Benchmarks lie if you don't read the fine print. Each axis, what it means, and where it bites:

Artificial Analysis Intelligence Index ↑ better
Composite 0–100 score averaging nine hard evaluations across agents, coding, general capability, and scientific reasoning.
Versioned composite (v4.1.1, 9 evals). Only comparable within one index version — pin the date.
MMLU-Pro ↑ better
Multiple-choice accuracy across 14 academic and professional domains — a harder MMLU successor with 10 answer options.
GPQA Diamond ↑ better
Accuracy on 198 graduate-level physics/chemistry/biology questions PhD experts answer ~65% of the time.
Nearing saturation (~94% at the top), compressing the high end.
Humanity's Last Exam ↑ better
~3,000 expert-crowdsourced, deliberately frontier-breaking questions across math, sciences, and the humanities.
Heavily reasoning-effort and tool-use dependent. Values are no-tools, max-reasoning.
SWE-bench Verified ↑ better
Percent of 500 human-validated real GitHub issues resolved with a patch that passes the repo’s hidden tests.
Swings 5–15 pts on the same model by harness (single-shot vs. agentic loop). Compare like with like.
LiveCodeBench ↑ better
Pass rate on competitive-programming problems released after a model’s training cutoff — contamination-free coding.
AIME 2025 ↑ better
Fraction of the 30 problems from the 2025 American Invitational Mathematics Examination solved correctly.
Saturated — multiple frontier models hit 100%. Stops discriminating at the top.
LMArena Elo ↑ better
Bradley-Terry rating from millions of blind, pairwise human preference votes on head-to-head chat responses.
Relative + re-anchoring as models enter. Style-control on/off shifts rankings.
Output speed ↑ better
Median output throughput in tokens generated per second during a single request.
A property of the provider endpoint, not the weights. The same model runs 5–20× faster on Cerebras/Groq.
Price (blended) ↓ better
Cost per million tokens, blended by Artificial Analysis at a 7:2:1 cache:input:output ratio.
Lower is better. The blend ratio is an editorial choice — reconcile against provider pages.
Context window ↑ better
Maximum input tokens a model can attend to in a single request, as advertised.
Advertised ≠ effective. RULER-style tests put usable context at ~50–65% of the label.
VRAM at Q4 ↓ better
Approximate GPU memory to run the model locally at Q4_K_M quantization (weights + modest context).
Lower is better. Open-weight only. Estimate ≈ params(B) × 0.55 + overhead.

The full matrix · 42 models

Every model, every metric. Cells brighten toward the leader in each column; the leader is ringed. Each number links to its source.

ModelAA IndexGPQASWE-benchArenaSpeedPriceContextVRAM
Claude Opus 5cloud6393.7%96%149352 tok/s$3.85/M1M
Claude Fable 5cloud6295%150667 tok/s$7.70/M1M
GPT-5.6 Solcloud6194.1%148269 tok/s$4.35/M1M
Grok 4.6cloud6194.9%146458 tok/s$1.35/M500K
Kimi K3open6093.5%38 tok/s$2.31/M1M1529 GB
Qwen3.8 Maxopen5892.6%149145 tok/s$1.18/M1M1345 GB
Claude Opus 4.8cloud5793.6%88.6%148358 tok/s$3.85/M1M
GPT-5.6 Terracloud57149 tok/s$1.74/M1M
Gemini 3.7 Flashcloud5694.5%1490368 tok/s$0.58/M1M
GPT-5.5cloud5693.5%88.7%148293 tok/s$4.35/M922K
Grok 4.5cloud5660 tok/s$1.21/M500K
Claude Sonnet 5cloud5572.7%82 tok/s$1.54/M1M
Claude Opus 4.7cloud5594.2%87.6%150155 tok/s$3.85/M1M
DeepSeek V4 Pro 0813open5379 tok/s$0.69/M1M908 GB
GLM-5.2open53141 tok/s$0.86/M1M
GPT-5.6 Lunacloud52202 tok/s$0.17/M1M
Gemini 3.6 Flashcloud521485238 tok/s$1.16/M1M
Gemini 3.5 Flashcloud521476191 tok/s$1.31/M1M
DeepSeek V4 Flash 0731open5291%141 tok/s$0.06/M1M167 GB
Qwen3.8-27Bopen5289.2%262K15 GB
Gemini 3.1 Procloud4894.1%80.6%1487136 tok/s$1.74/M1M
Claude Sonnet 4.6cloud4880.2%144255 tok/s$2.31/M1M
Qwen3.7 Maxcloud4792.4%1460210 tok/s$1.60/M1M
GPT-5.3 Codexcloud46145 tok/s$1.87/M400K
MiniMax-M3cloud4592.9%94 tok/s$0.22/M1M
Kimi K2.6open4590.5%80.2%142254 tok/s$0.70/M256K
DeepSeek V4 Proopen4590.1%80.6%78 tok/s$0.18/M1M517 GB
Muse Sparkcloud441488262K
Kimi K2.7 Codeopen4344 tok/s$0.72/M256K
MiMo-V2.5-Proopen43142577 tok/s$0.18/M1M
GLM-5.1open4186.2%146772 tok/s$0.90/M200K
GPT-5.4 minicloud41184 tok/s$0.65/M400K
DeepSeek V4 Flash (Preview)open4088.1%79%122 tok/s$0.06/M1M95 GB
Qwen3.7 Pluscloud3957 tok/s$0.27/M1M
Grok 4.3cloud38129 tok/s$0.64/M1M
Gemini 3.5 Flash-Litecloud37397 tok/s$0.33/M1M
Gemma 3 27Bopen41.4%131K16 GB
gpt-oss-120bopen80.1%131K63 GB
gpt-oss-20bopen71.5%131K13 GB
Mistral Small 3.2 24Bopen46.1%131K14 GB
Phi-4open56.1%16K9 GB
Qwen3-32Bopen41K20 GB

Method

A hand-curated snapshot, verified 2026-06-06. Standardized cross-model metrics (Intelligence Index, output speed, blended price, context) come from Artificial Analysis; human-preference Elo from LMArena; SWE-bench Verified, GPQA, HLE, MMLU-Pro and LiveCodeBench from vendor model cards and announcement posts; open-weight parameter counts from the HuggingFace safetensors index.

Every number on this page carries a source link — hover a node, open the detail card, or click any cell. Where a model hasn't published a number, the cell reads "—" rather than a guess. The model writes none of these figures; the catalog does.

Snapshot, not a live feed. Benchmarks saturate, leaderboards re-anchor, and the frontier moves weekly — re-verify against the linked sources before you cut a PO against them.