Skip to content
MCS · Models · 42

Model Ratings.

The tracked cohort, scored and tiered.

42 models ranked against the frontier intelligence recipe · every axis sourced.

Scores are relative to the tracked cohort — swap the recipe or the field and the ranks move. The weights are an editorial call; the axis values underneath are sourced, cell by cell.

Snapshot · verified 1w ago · 2026-08-17

Recipe

Raw reasoning, knowledge, and agentic coding — price-blind.

The board · ranked · Frontier intelligence

Each cell brightens toward the cohort leader and is ringed when it holds it; dimmed rows are sparse (low coverage) — read their composite with care.

  1. 1
    Claude Fable 5
    Anthropic
    67% · mediumS98
    1. AA Index62
    2. GPQA
    3. HLE55.5%
    4. SWE-bench95%
    5. Arena1506
    6. LiveCode

    Strongest: aa_index · Weakest: swebench

  2. 2
    Claude Opus 5
    Anthropic
    83% · highS97
    1. AA Index63
    2. GPQA93.7%
    3. HLE54.9%
    4. SWE-bench96%
    5. Arena1493
    6. LiveCode

    Strongest: aa_index · Weakest: arena_elo

  3. 3
    GPT-5.6 Sol
    OpenAI
    67% · mediumS88
    1. AA Index61
    2. GPQA94.1%
    3. HLE47.2%
    4. SWE-bench
    5. Arena1482
    6. LiveCode

    Strongest: aa_index · Weakest: arena_elo

  4. 4
    Kimi K3
    Moonshot AI
    50% · mediumS87
    1. AA Index60
    2. GPQA93.5%
    3. HLE43.5%
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: hle

  5. 5
    Grok 4.6
    xAI
    50% · mediumS86
    1. AA Index61
    2. GPQA94.9%
    3. HLE
    4. SWE-bench
    5. Arena1464
    6. LiveCode

    Strongest: aa_index · Weakest: arena_elo

  6. 6
    Qwen3.8 Max
    Alibaba (Qwen)
    67% · mediumA83
    1. AA Index58
    2. GPQA92.6%
    3. HLE43.6%
    4. SWE-bench
    5. Arena1491
    6. LiveCode

    Strongest: aa_index · Weakest: hle

  7. 7
    Gemini 3.7 Flash
    Google DeepMind
    50% · mediumA83
    1. AA Index56
    2. GPQA94.5%
    3. HLE
    4. SWE-bench
    5. Arena1490
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  8. 8
    Claude Opus 4.7
    Anthropic
    83% · highA79
    1. AA Index55
    2. GPQA94.2%
    3. HLE46.9%
    4. SWE-bench87.6%
    5. Arena1501
    6. LiveCode

    Strongest: aa_index · Weakest: swebench

  9. 9
    Claude Opus 4.8
    Anthropic
    83% · highA79
    1. AA Index57
    2. GPQA93.6%
    3. HLE45.7%
    4. SWE-bench88.6%
    5. Arena1483
    6. LiveCode

    Strongest: aa_index · Weakest: swebench

  10. 10
    GPT-5.5
    OpenAI
    83% · highA77
    1. AA Index56
    2. GPQA93.5%
    3. HLE44.3%
    4. SWE-bench88.7%
    5. Arena1482
    6. LiveCode

    Strongest: aa_index · Weakest: swebench

  11. 11
    GPT-5.6 Terra
    OpenAI
    17% · lowA77
    1. AA Index57
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  12. 12
    Grok 4.5
    xAI
    17% · lowA73
    1. AA Index56
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  13. 13
    DeepSeek V4 Flash 0731
    DeepSeek
    50% · mediumB68
    1. AA Index52
    2. GPQA91%
    3. HLE37%
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  14. 14
    Qwen3.8-27B
    Alibaba (Qwen)
    67% · mediumB66
    1. AA Index52
    2. GPQA89.2%
    3. HLE30.8%
    4. SWE-bench
    5. Arena
    6. LiveCode90.3%

    Strongest: aa_index · Weakest: hle

  15. 15
    DeepSeek V4 Pro 0813
    DeepSeek
    33% · lowB65
    1. AA Index53
    2. GPQA
    3. HLE42.7%
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  16. 16
    Gemini 3.6 Flash
    Google DeepMind
    33% · lowB63
    1. AA Index52
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena1485
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  17. 17
    Gemini 3.1 Pro
    Google DeepMind
    83% · highB62
    1. AA Index48
    2. GPQA94.1%
    3. HLE44.7%
    4. SWE-bench80.6%
    5. Arena1487
    6. LiveCode

    Strongest: gpqa · Weakest: swebench

  18. 18
    GLM-5.2
    Z.AI (Zhipu)
    17% · lowB62
    1. AA Index53
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  19. 19
    Gemini 3.5 Flash
    Google DeepMind
    50% · mediumB61
    1. AA Index52
    2. GPQA
    3. HLE40.2%
    4. SWE-bench
    5. Arena1476
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  20. 20
    GPT-5.6 Luna
    OpenAI
    17% · lowB58
    1. AA Index52
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  21. 21
    DeepSeek V4 Pro
    DeepSeek
    83% · highB56
    1. AA Index45
    2. GPQA90.1%
    3. HLE37.7%
    4. SWE-bench80.6%
    5. Arena
    6. LiveCode93.5%

    Strongest: gpqa · Weakest: aa_index

  22. 22
    MiniMax-M3
    MiniMax
    33% · lowB55
    1. AA Index45
    2. GPQA92.9%
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: gpqa · Weakest: aa_index

  23. 23
    Claude Sonnet 5
    Anthropic
    50% · mediumC53
    1. AA Index55
    2. GPQA
    3. HLE43.2%
    4. SWE-bench72.7%
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: swebench

  24. 24
    Qwen3.7 Max
    Alibaba (Qwen)
    83% · highC51
    1. AA Index47
    2. GPQA92.4%
    3. HLE38.2%
    4. SWE-bench
    5. Arena1460
    6. LiveCode78.2%

    Strongest: gpqa · Weakest: livecodebench

  25. 25
    Kimi K2.6
    Moonshot AI
    100% · highC46
    1. AA Index45
    2. GPQA90.5%
    3. HLE34.7%
    4. SWE-bench80.2%
    5. Arena1422
    6. LiveCode89.6%

    Strongest: gpqa · Weakest: arena_elo

  26. 26
    DeepSeek V4 Flash (Preview)
    DeepSeek
    83% · highC45
    1. AA Index40
    2. GPQA88.1%
    3. HLE34.8%
    4. SWE-bench79%
    5. Arena
    6. LiveCode91.6%

    Strongest: gpqa · Weakest: aa_index

  27. 27
    GLM-5.1
    Z.AI (Zhipu)
    67% · mediumC44
    1. AA Index41
    2. GPQA86.2%
    3. HLE31%
    4. SWE-bench
    5. Arena1467
    6. LiveCode

    Strongest: gpqa · Weakest: aa_index

  28. 28
    gpt-oss-120b
    OpenAI
    33% · lowC44
    1. AA Index
    2. GPQA80.1%
    3. HLE14.9%
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: gpqa · Weakest: hle

  29. 29
    Muse Spark
    Meta
    33% · lowC42
    1. AA Index44
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena1488
    6. LiveCode

    Strongest: arena_elo · Weakest: aa_index

  30. 30
    Claude Sonnet 4.6
    Anthropic
    50% · mediumD36
    1. AA Index48
    2. GPQA
    3. HLE
    4. SWE-bench80.2%
    5. Arena1442
    6. LiveCode

    Strongest: aa_index · Weakest: arena_elo

  31. 31
    GPT-5.3 Codex
    OpenAI
    17% · lowD35
    1. AA Index46
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  32. 32
    gpt-oss-20b
    OpenAI
    33% · lowD31
    1. AA Index
    2. GPQA71.5%
    3. HLE10.9%
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: gpqa · Weakest: hle

  33. 33
    Phi-4
    Microsoft
    17% · lowD27
    1. AA Index
    2. GPQA56.1%
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: gpqa · Weakest: gpqa

  34. 34
    Kimi K2.7 Code
    Moonshot AI
    17% · lowD23
    1. AA Index43
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  35. 35
    MiMo-V2.5-Pro
    Xiaomi
    33% · lowD18
    1. AA Index43
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena1425
    6. LiveCode

    Strongest: aa_index · Weakest: arena_elo

  36. 36
    GPT-5.4 mini
    OpenAI
    17% · lowD15
    1. AA Index41
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  37. 37
    Mistral Small 3.2 24B
    Mistral AI
    17% · lowD9
    1. AA Index
    2. GPQA46.1%
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: gpqa · Weakest: gpqa

  38. 38
    Qwen3.7 Plus
    Alibaba (Qwen)
    17% · lowD8
    1. AA Index39
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  39. 39
    Grok 4.3
    xAI
    17% · lowD4
    1. AA Index38
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  40. 40
    Gemini 3.5 Flash-Lite
    Google DeepMind
    17% · lowD0
    1. AA Index37
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: aa_index · Weakest: aa_index

  41. 41
    Gemma 3 27B
    Google
    17% · lowD0
    1. AA Index
    2. GPQA41.4%
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: gpqa · Weakest: gpqa

  42. 42
    Qwen3-32B
    Alibaba (Qwen)
    0% · lowD0
    1. AA Index
    2. GPQA
    3. HLE
    4. SWE-bench
    5. Arena
    6. LiveCode

    Strongest: · Weakest:

Composite = Σ(goodness · weight) over the criteria a model actually carries, renormalized so a sparse row scores on its axes — not against zeros. Tiers map the 0–100 composite onto S / A / B / C / D. Re-verify against the linked sources before you cut a PO.

Model Ratings — MadCoolStuff · MadCoolStuff