Skip to content
MCS · Models · 42

Model Ratings.

The tracked cohort, scored and tiered.

42 models ranked against the coding muscle recipe · every axis sourced.

Scores are relative to the tracked cohort — swap the recipe or the field and the ranks move. The weights are an editorial call; the axis values underneath are sourced, cell by cell.

Snapshot · verified 1w ago · 2026-08-17

Recipe

Ship-real-PRs ability — agentic + competitive coding weighted hardest.

The board · ranked · Coding muscle

Each cell brightens toward the cohort leader and is ringed when it holds it; dimmed rows are sparse (low coverage) — read their composite with care.

  1. 1
    Claude Opus 5
    Anthropic
    75% · highS99
    1. SWE-bench96%
    2. LiveCode
    3. AA Index63
    4. GPQA93.7%

    Strongest: swebench · Weakest: gpqa

  2. 2
    Claude Fable 5
    Anthropic
    50% · mediumS96
    1. SWE-bench95%
    2. LiveCode
    3. AA Index62
    4. GPQA

    Strongest: swebench · Weakest: swebench

  3. 3
    Grok 4.6
    xAI
    50% · mediumS96
    1. SWE-bench
    2. LiveCode
    3. AA Index61
    4. GPQA94.9%

    Strongest: aa_index · Weakest: aa_index

  4. 4
    GPT-5.6 Sol
    OpenAI
    50% · mediumS95
    1. SWE-bench
    2. LiveCode
    3. AA Index61
    4. GPQA94.1%

    Strongest: aa_index · Weakest: aa_index

  5. 5
    Kimi K3
    Moonshot AI
    50% · mediumS92
    1. SWE-bench
    2. LiveCode
    3. AA Index60
    4. GPQA93.5%

    Strongest: aa_index · Weakest: aa_index

  6. 6
    Qwen3.8 Max
    Alibaba (Qwen)
    50% · mediumS87
    1. SWE-bench
    2. LiveCode
    3. AA Index58
    4. GPQA92.6%

    Strongest: aa_index · Weakest: aa_index

  7. 7
    Gemini 3.7 Flash
    Google DeepMind
    50% · mediumA85
    1. SWE-bench
    2. LiveCode
    3. AA Index56
    4. GPQA94.5%

    Strongest: gpqa · Weakest: aa_index

  8. 8
    Claude Opus 4.8
    Anthropic
    75% · highA77
    1. SWE-bench88.6%
    2. LiveCode
    3. AA Index57
    4. GPQA93.6%

    Strongest: swebench · Weakest: swebench

  9. 9
    GPT-5.6 Terra
    OpenAI
    25% · lowA77
    1. SWE-bench
    2. LiveCode
    3. AA Index57
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  10. 10
    GPT-5.5
    OpenAI
    75% · highA76
    1. SWE-bench88.7%
    2. LiveCode
    3. AA Index56
    4. GPQA93.5%

    Strongest: swebench · Weakest: swebench

  11. 11
    Qwen3.8-27B
    Alibaba (Qwen)
    75% · highA75
    1. SWE-bench
    2. LiveCode90.3%
    3. AA Index52
    4. GPQA89.2%

    Strongest: livecodebench · Weakest: aa_index

  12. 12
    Claude Opus 4.7
    Anthropic
    75% · highA73
    1. SWE-bench87.6%
    2. LiveCode
    3. AA Index55
    4. GPQA94.2%

    Strongest: swebench · Weakest: swebench

  13. 13
    DeepSeek V4 Flash 0731
    DeepSeek
    50% · mediumA73
    1. SWE-bench
    2. LiveCode
    3. AA Index52
    4. GPQA91%

    Strongest: gpqa · Weakest: aa_index

  14. 14
    Grok 4.5
    xAI
    25% · lowA73
    1. SWE-bench
    2. LiveCode
    3. AA Index56
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  15. 15
    gpt-oss-120b
    OpenAI
    25% · lowA72
    1. SWE-bench
    2. LiveCode
    3. AA Index
    4. GPQA80.1%

    Strongest: gpqa · Weakest: gpqa

  16. 16
    DeepSeek V4 Pro
    DeepSeek
    100% · highB62
    1. SWE-bench80.6%
    2. LiveCode93.5%
    3. AA Index45
    4. GPQA90.1%

    Strongest: livecodebench · Weakest: aa_index

  17. 17
    DeepSeek V4 Pro 0813
    DeepSeek
    25% · lowB62
    1. SWE-bench
    2. LiveCode
    3. AA Index53
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  18. 18
    GLM-5.2
    Z.AI (Zhipu)
    25% · lowB62
    1. SWE-bench
    2. LiveCode
    3. AA Index53
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  19. 19
    MiniMax-M3
    MiniMax
    50% · mediumB60
    1. SWE-bench
    2. LiveCode
    3. AA Index45
    4. GPQA92.9%

    Strongest: gpqa · Weakest: aa_index

  20. 20
    Gemini 3.5 Flash
    Google DeepMind
    25% · lowB58
    1. SWE-bench
    2. LiveCode
    3. AA Index52
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  21. 21
    Gemini 3.6 Flash
    Google DeepMind
    25% · lowB58
    1. SWE-bench
    2. LiveCode
    3. AA Index52
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  22. 22
    GPT-5.6 Luna
    OpenAI
    25% · lowB58
    1. SWE-bench
    2. LiveCode
    3. AA Index52
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  23. 23
    gpt-oss-20b
    OpenAI
    25% · lowB56
    1. SWE-bench
    2. LiveCode
    3. AA Index
    4. GPQA71.5%

    Strongest: gpqa · Weakest: gpqa

  24. 24
    Kimi K2.6
    Moonshot AI
    100% · highC54
    1. SWE-bench80.2%
    2. LiveCode89.6%
    3. AA Index45
    4. GPQA90.5%

    Strongest: livecodebench · Weakest: aa_index

  25. 25
    DeepSeek V4 Flash (Preview)
    DeepSeek
    100% · highC52
    1. SWE-bench79%
    2. LiveCode91.6%
    3. AA Index40
    4. GPQA88.1%

    Strongest: livecodebench · Weakest: aa_index

  26. 26
    Gemini 3.1 Pro
    Google DeepMind
    75% · highC51
    1. SWE-bench80.6%
    2. LiveCode
    3. AA Index48
    4. GPQA94.1%

    Strongest: gpqa · Weakest: swebench

  27. 27
    GLM-5.1
    Z.AI (Zhipu)
    50% · mediumC46
    1. SWE-bench
    2. LiveCode
    3. AA Index41
    4. GPQA86.2%

    Strongest: gpqa · Weakest: aa_index

  28. 28
    Claude Sonnet 4.6
    Anthropic
    50% · mediumD36
    1. SWE-bench80.2%
    2. LiveCode
    3. AA Index48
    4. GPQA

    Strongest: swebench · Weakest: swebench

  29. 29
    Qwen3.7 Max
    Alibaba (Qwen)
    75% · highD35
    1. SWE-bench
    2. LiveCode78.2%
    3. AA Index47
    4. GPQA92.4%

    Strongest: gpqa · Weakest: livecodebench

  30. 30
    GPT-5.3 Codex
    OpenAI
    25% · lowD35
    1. SWE-bench
    2. LiveCode
    3. AA Index46
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  31. 31
    Phi-4
    Microsoft
    25% · lowD27
    1. SWE-bench
    2. LiveCode
    3. AA Index
    4. GPQA56.1%

    Strongest: gpqa · Weakest: gpqa

  32. 32
    Muse Spark
    Meta
    25% · lowD27
    1. SWE-bench
    2. LiveCode
    3. AA Index44
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  33. 33
    Claude Sonnet 5
    Anthropic
    50% · mediumD26
    1. SWE-bench72.7%
    2. LiveCode
    3. AA Index55
    4. GPQA

    Strongest: aa_index · Weakest: swebench

  34. 34
    Kimi K2.7 Code
    Moonshot AI
    25% · lowD23
    1. SWE-bench
    2. LiveCode
    3. AA Index43
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  35. 35
    MiMo-V2.5-Pro
    Xiaomi
    25% · lowD23
    1. SWE-bench
    2. LiveCode
    3. AA Index43
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  36. 36
    GPT-5.4 mini
    OpenAI
    25% · lowD15
    1. SWE-bench
    2. LiveCode
    3. AA Index41
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  37. 37
    Mistral Small 3.2 24B
    Mistral AI
    25% · lowD9
    1. SWE-bench
    2. LiveCode
    3. AA Index
    4. GPQA46.1%

    Strongest: gpqa · Weakest: gpqa

  38. 38
    Qwen3.7 Plus
    Alibaba (Qwen)
    25% · lowD8
    1. SWE-bench
    2. LiveCode
    3. AA Index39
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  39. 39
    Grok 4.3
    xAI
    25% · lowD4
    1. SWE-bench
    2. LiveCode
    3. AA Index38
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  40. 40
    Gemini 3.5 Flash-Lite
    Google DeepMind
    25% · lowD0
    1. SWE-bench
    2. LiveCode
    3. AA Index37
    4. GPQA

    Strongest: aa_index · Weakest: aa_index

  41. 41
    Gemma 3 27B
    Google
    25% · lowD0
    1. SWE-bench
    2. LiveCode
    3. AA Index
    4. GPQA41.4%

    Strongest: gpqa · Weakest: gpqa

  42. 42
    Qwen3-32B
    Alibaba (Qwen)
    0% · lowD0
    1. SWE-bench
    2. LiveCode
    3. AA Index
    4. GPQA

    Strongest: · Weakest:

Composite = Σ(goodness · weight) over the criteria a model actually carries, renormalized so a sparse row scores on its axes — not against zeros. Tiers map the 0–100 composite onto S / A / B / C / D. Re-verify against the linked sources before you cut a PO.

Model Ratings — MadCoolStuff · MadCoolStuff