Model Ratings.
The tracked cohort, scored and tiered.
42 models ranked against the frontier intelligence recipe · every axis sourced.
Scores are relative to the tracked cohort — swap the recipe or the field and the ranks move. The weights are an editorial call; the axis values underneath are sourced, cell by cell.
Snapshot · verified 1w ago · 2026-08-17
Recipe
Raw reasoning, knowledge, and agentic coding — price-blind.
The board · ranked · Frontier intelligence
Each cell brightens toward the cohort leader and is ringed when it holds it; dimmed rows are sparse (low coverage) — read their composite with care.
- 1Claude Fable 5Anthropic67% · mediumS98
- AA Index62
- GPQA—
- HLE55.5%
- SWE-bench95%
- Arena1506
- LiveCode—
Strongest: aa_index · Weakest: swebench
- 2Claude Opus 5Anthropic83% · highS97
Strongest: aa_index · Weakest: arena_elo
- 3GPT-5.6 SolOpenAI67% · mediumS88
- AA Index61
- GPQA94.1%
- HLE47.2%
- SWE-bench—
- Arena1482
- LiveCode—
Strongest: aa_index · Weakest: arena_elo
- 4Kimi K3Moonshot AI50% · mediumS87
- AA Index60
- GPQA93.5%
- HLE43.5%
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: hle
- 5Grok 4.6xAI50% · mediumS86
- AA Index61
- GPQA94.9%
- HLE—
- SWE-bench—
- Arena1464
- LiveCode—
Strongest: aa_index · Weakest: arena_elo
- 6Qwen3.8 MaxAlibaba (Qwen)67% · mediumA83
- AA Index58
- GPQA92.6%
- HLE43.6%
- SWE-bench—
- Arena1491
- LiveCode—
Strongest: aa_index · Weakest: hle
- 7Gemini 3.7 FlashGoogle DeepMind50% · mediumA83
- AA Index56
- GPQA94.5%
- HLE—
- SWE-bench—
- Arena1490
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 8Claude Opus 4.7Anthropic83% · highA79
Strongest: aa_index · Weakest: swebench
- 9Claude Opus 4.8Anthropic83% · highA79
Strongest: aa_index · Weakest: swebench
- 10GPT-5.5OpenAI83% · highA77
Strongest: aa_index · Weakest: swebench
- 11GPT-5.6 TerraOpenAI17% · lowA77
- AA Index57
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 12Grok 4.5xAI17% · lowA73
- AA Index56
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 13DeepSeek V4 Flash 0731DeepSeek50% · mediumB68
- AA Index52
- GPQA91%
- HLE37%
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 14Qwen3.8-27BAlibaba (Qwen)67% · mediumB66
- AA Index52
- GPQA89.2%
- HLE30.8%
- SWE-bench—
- Arena—
- LiveCode90.3%
Strongest: aa_index · Weakest: hle
- 15DeepSeek V4 Pro 0813DeepSeek33% · lowB65
- AA Index53
- GPQA—
- HLE42.7%
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 16Gemini 3.6 FlashGoogle DeepMind33% · lowB63
- AA Index52
- GPQA—
- HLE—
- SWE-bench—
- Arena1485
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 17Gemini 3.1 ProGoogle DeepMind83% · highB62
Strongest: gpqa · Weakest: swebench
- 18GLM-5.2Z.AI (Zhipu)17% · lowB62
- AA Index53
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 19Gemini 3.5 FlashGoogle DeepMind50% · mediumB61
- AA Index52
- GPQA—
- HLE40.2%
- SWE-bench—
- Arena1476
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 20GPT-5.6 LunaOpenAI17% · lowB58
- AA Index52
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 21DeepSeek V4 ProDeepSeek83% · highB56
Strongest: gpqa · Weakest: aa_index
- 22MiniMax-M3MiniMax33% · lowB55
- AA Index45
- GPQA92.9%
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: gpqa · Weakest: aa_index
- 23Claude Sonnet 5Anthropic50% · mediumC53
- AA Index55
- GPQA—
- HLE43.2%
- SWE-bench72.7%
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: swebench
- 24Qwen3.7 MaxAlibaba (Qwen)83% · highC51
Strongest: gpqa · Weakest: livecodebench
- 25Kimi K2.6Moonshot AI100% · highC46
Strongest: gpqa · Weakest: arena_elo
- 26DeepSeek V4 Flash (Preview)DeepSeek83% · highC45
Strongest: gpqa · Weakest: aa_index
- 27GLM-5.1Z.AI (Zhipu)67% · mediumC44
- AA Index41
- GPQA86.2%
- HLE31%
- SWE-bench—
- Arena1467
- LiveCode—
Strongest: gpqa · Weakest: aa_index
- 28gpt-oss-120bOpenAI33% · lowC44
Strongest: gpqa · Weakest: hle
- 29Muse SparkMeta33% · lowC42
- AA Index44
- GPQA—
- HLE—
- SWE-bench—
- Arena1488
- LiveCode—
Strongest: arena_elo · Weakest: aa_index
- 30Claude Sonnet 4.6Anthropic50% · mediumD36
- AA Index48
- GPQA—
- HLE—
- SWE-bench80.2%
- Arena1442
- LiveCode—
Strongest: aa_index · Weakest: arena_elo
- 31GPT-5.3 CodexOpenAI17% · lowD35
- AA Index46
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 32gpt-oss-20bOpenAI33% · lowD31
Strongest: gpqa · Weakest: hle
- 33Phi-4Microsoft17% · lowD27
- AA Index—
- GPQA56.1%
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: gpqa · Weakest: gpqa
- 34Kimi K2.7 CodeMoonshot AI17% · lowD23
- AA Index43
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 35MiMo-V2.5-ProXiaomi33% · lowD18
- AA Index43
- GPQA—
- HLE—
- SWE-bench—
- Arena1425
- LiveCode—
Strongest: aa_index · Weakest: arena_elo
- 36GPT-5.4 miniOpenAI17% · lowD15
- AA Index41
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 37Mistral Small 3.2 24BMistral AI17% · lowD9
- AA Index—
- GPQA46.1%
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: gpqa · Weakest: gpqa
- 38Qwen3.7 PlusAlibaba (Qwen)17% · lowD8
- AA Index39
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 39Grok 4.3xAI17% · lowD4
- AA Index38
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 40Gemini 3.5 Flash-LiteGoogle DeepMind17% · lowD0
- AA Index37
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: aa_index · Weakest: aa_index
- 41Gemma 3 27BGoogle17% · lowD0
- AA Index—
- GPQA41.4%
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: gpqa · Weakest: gpqa
- 42Qwen3-32BAlibaba (Qwen)0% · lowD0
- AA Index—
- GPQA—
- HLE—
- SWE-bench—
- Arena—
- LiveCode—
Strongest: — · Weakest: —