Skip to content
MCS · Models · 42

Model Ratings.

The tracked cohort, scored and tiered.

42 models ranked against the runs local recipe · every axis sourced.

Scores are relative to the tracked cohort — swap the recipe or the field and the ranks move. The weights are an editorial call; the axis values underneath are sourced, cell by cell.

Snapshot · verified 1w ago · 2026-08-17

Recipe

Open-weight, fits your GPU — VRAM and context weighted alongside smarts.

The board · ranked · Runs local

Each cell brightens toward the cohort leader and is ringed when it holds it; dimmed rows are sparse (low coverage) — read their composite with care.

  1. 1
    Claude Opus 5
    Anthropic
    60% · mediumS100
    1. VRAM
    2. AA Index63
    3. SWE-bench96%
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  2. 2
    Claude Fable 5
    Anthropic
    60% · mediumS97
    1. VRAM
    2. AA Index62
    3. SWE-bench95%
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: swebench

  3. 3
    GPT-5.6 Sol
    OpenAI
    40% · mediumS95
    1. VRAM
    2. AA Index61
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  4. 4
    GPT-5.6 Terra
    OpenAI
    40% · mediumA85
    1. VRAM
    2. AA Index57
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  5. 5
    Gemini 3.7 Flash
    Google DeepMind
    40% · mediumA82
    1. VRAM
    2. AA Index56
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  6. 6
    DeepSeek V4 Flash 0731
    DeepSeek
    60% · mediumA80
    1. VRAM167 GB
    2. AA Index52
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: vram_q4 · Weakest: aa_index

  7. 7
    Claude Opus 4.8
    Anthropic
    60% · mediumA79
    1. VRAM
    2. AA Index57
    3. SWE-bench88.6%
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: swebench

  8. 8
    Grok 4.6
    xAI
    40% · mediumA78
    1. VRAM
    2. AA Index61
    3. SWE-bench
    4. LiveCode
    5. Context500K

    Strongest: aa_index · Weakest: context

  9. 9
    GPT-5.5
    OpenAI
    60% · mediumA76
    1. VRAM
    2. AA Index56
    3. SWE-bench88.7%
    4. LiveCode
    5. Context922K

    Strongest: aa_index · Weakest: swebench

  10. 10
    gpt-oss-20b
    OpenAI
    40% · mediumA75
    1. VRAM13 GB
    2. AA Index
    3. SWE-bench
    4. LiveCode
    5. Context131K

    Strongest: vram_q4 · Weakest: context

  11. 11
    Mistral Small 3.2 24B
    Mistral AI
    40% · mediumA75
    1. VRAM14 GB
    2. AA Index
    3. SWE-bench
    4. LiveCode
    5. Context131K

    Strongest: vram_q4 · Weakest: context

  12. 12
    Gemma 3 27B
    Google
    40% · mediumA74
    1. VRAM16 GB
    2. AA Index
    3. SWE-bench
    4. LiveCode
    5. Context131K

    Strongest: vram_q4 · Weakest: context

  13. 13
    GLM-5.2
    Z.AI (Zhipu)
    40% · mediumA74
    1. VRAM
    2. AA Index53
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  14. 14
    Claude Opus 4.7
    Anthropic
    60% · mediumA74
    1. VRAM
    2. AA Index55
    3. SWE-bench87.6%
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: swebench

  15. 15
    Qwen3.8-27B
    Alibaba (Qwen)
    80% · highA72
    1. VRAM15 GB
    2. AA Index52
    3. SWE-bench
    4. LiveCode90.3%
    5. Context262K

    Strongest: vram_q4 · Weakest: context

  16. 16
    gpt-oss-120b
    OpenAI
    40% · mediumA72
    1. VRAM63 GB
    2. AA Index
    3. SWE-bench
    4. LiveCode
    5. Context131K

    Strongest: vram_q4 · Weakest: context

  17. 17
    Gemini 3.5 Flash
    Google DeepMind
    40% · mediumA72
    1. VRAM
    2. AA Index52
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  18. 18
    Gemini 3.6 Flash
    Google DeepMind
    40% · mediumA72
    1. VRAM
    2. AA Index52
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  19. 19
    GPT-5.6 Luna
    OpenAI
    40% · mediumA72
    1. VRAM
    2. AA Index52
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: aa_index

  20. 20
    Qwen3-32B
    Alibaba (Qwen)
    40% · mediumA72
    1. VRAM20 GB
    2. AA Index
    3. SWE-bench
    4. LiveCode
    5. Context41K

    Strongest: vram_q4 · Weakest: context

  21. 21
    Phi-4
    Microsoft
    40% · mediumA71
    1. VRAM9 GB
    2. AA Index
    3. SWE-bench
    4. LiveCode
    5. Context16K

    Strongest: vram_q4 · Weakest: context

  22. 22
    Grok 4.5
    xAI
    40% · mediumB65
    1. VRAM
    2. AA Index56
    3. SWE-bench
    4. LiveCode
    5. Context500K

    Strongest: aa_index · Weakest: context

  23. 23
    DeepSeek V4 Flash (Preview)
    DeepSeek
    100% · highB61
    1. VRAM95 GB
    2. AA Index40
    3. SWE-bench79%
    4. LiveCode91.6%
    5. Context1M

    Strongest: vram_q4 · Weakest: aa_index

  24. 24
    DeepSeek V4 Pro
    DeepSeek
    100% · highB60
    1. VRAM517 GB
    2. AA Index45
    3. SWE-bench80.6%
    4. LiveCode93.5%
    5. Context1M

    Strongest: vram_q4 · Weakest: aa_index

  25. 25
    DeepSeek V4 Pro 0813
    DeepSeek
    60% · mediumB59
    1. VRAM908 GB
    2. AA Index53
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: vram_q4

  26. 26
    MiniMax-M3
    MiniMax
    40% · mediumC54
    1. VRAM
    2. AA Index45
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: aa_index

  27. 27
    Qwen3.8 Max
    Alibaba (Qwen)
    60% · mediumC53
    1. VRAM1345 GB
    2. AA Index58
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: vram_q4

  28. 28
    Gemini 3.1 Pro
    Google DeepMind
    60% · mediumC52
    1. VRAM
    2. AA Index48
    3. SWE-bench80.6%
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: swebench

  29. 29
    Claude Sonnet 5
    Anthropic
    60% · mediumC51
    1. VRAM
    2. AA Index55
    3. SWE-bench72.7%
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: swebench

  30. 30
    Claude Sonnet 4.6
    Anthropic
    60% · mediumC51
    1. VRAM
    2. AA Index48
    3. SWE-bench80.2%
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: swebench

  31. 31
    Kimi K3
    Moonshot AI
    60% · mediumC50
    1. VRAM1529 GB
    2. AA Index60
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: aa_index · Weakest: vram_q4

  32. 32
    MiMo-V2.5-Pro
    Xiaomi
    40% · mediumC49
    1. VRAM
    2. AA Index43
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: aa_index

  33. 33
    Qwen3.7 Max
    Alibaba (Qwen)
    60% · mediumC42
    1. VRAM
    2. AA Index47
    3. SWE-bench
    4. LiveCode78.2%
    5. Context1M

    Strongest: context · Weakest: livecodebench

  34. 34
    Kimi K2.6
    Moonshot AI
    80% · highD39
    1. VRAM
    2. AA Index45
    3. SWE-bench80.2%
    4. LiveCode89.6%
    5. Context256K

    Strongest: livecodebench · Weakest: context

  35. 35
    Qwen3.7 Plus
    Alibaba (Qwen)
    40% · mediumD38
    1. VRAM
    2. AA Index39
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: aa_index

  36. 36
    GPT-5.3 Codex
    OpenAI
    40% · mediumD36
    1. VRAM
    2. AA Index46
    3. SWE-bench
    4. LiveCode
    5. Context400K

    Strongest: aa_index · Weakest: aa_index

  37. 37
    Grok 4.3
    xAI
    40% · mediumD36
    1. VRAM
    2. AA Index38
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: aa_index

  38. 38
    Gemini 3.5 Flash-Lite
    Google DeepMind
    40% · mediumD33
    1. VRAM
    2. AA Index37
    3. SWE-bench
    4. LiveCode
    5. Context1M

    Strongest: context · Weakest: aa_index

  39. 39
    Muse Spark
    Meta
    40% · mediumD26
    1. VRAM
    2. AA Index44
    3. SWE-bench
    4. LiveCode
    5. Context262K

    Strongest: aa_index · Weakest: context

  40. 40
    Kimi K2.7 Code
    Moonshot AI
    40% · mediumD24
    1. VRAM
    2. AA Index43
    3. SWE-bench
    4. LiveCode
    5. Context256K

    Strongest: aa_index · Weakest: aa_index

  41. 41
    GPT-5.4 mini
    OpenAI
    40% · mediumD23
    1. VRAM
    2. AA Index41
    3. SWE-bench
    4. LiveCode
    5. Context400K

    Strongest: context · Weakest: aa_index

  42. 42
    GLM-5.1
    Z.AI (Zhipu)
    40% · mediumD16
    1. VRAM
    2. AA Index41
    3. SWE-bench
    4. LiveCode
    5. Context200K

    Strongest: aa_index · Weakest: aa_index

Composite = Σ(goodness · weight) over the criteria a model actually carries, renormalized so a sparse row scores on its axes — not against zeros. Tiers map the 0–100 composite onto S / A / B / C / D. Re-verify against the linked sources before you cut a PO.

Model Ratings — MadCoolStuff · MadCoolStuff