howclose.to
The frontier · as of 2026-09-02

175 claims about 68 models. 53 have been checked by anyone else.

Every published benchmark result we could source for the current frontier and the open-weight field, tagged by who stands behind it. 30% carry an independent steward's check; the rest rest on the word of the party with an interest in the number. Neither is dismissed here — but they are never drawn the same way.What AI companies say their newest models can do — and whether anyone independent has actually tested it. Usually, nobody has. We show both, and we never draw them the same way.

Solid = an independent scorekeeper checked it. Hollow= the lab's own claim.

01 · Who checks

27 of 68 models have no independent result at all

Independent checks come from a short list of stewards: ARC Prize Foundation, Epoch AI, METR, SWE-bench maintainers. Where a model is absent from all of them, every number it has is its own. A leaderboard row that is merely submittedis not a check — SWE-bench's current entries sit unexamined, and are counted here as claims.Only a handful of independent groups actually re-run these tests: ARC Prize Foundation, Epoch AI, METR, SWE-bench maintainers. If a model is missing from all of them, every number about it came from the company that built it.

Claude Mythos 5.1

Anthropic2026-09closed (proprietary)

6 claimed · none independently checked

  • Terminal-Bench 4.060.9%Claude Code with --bare; max thinking; 10 trials per task; Mythos safeguards configurationsource
  • BioMysteryBench Human Solvable90.3%bash and file editor; restricted domains; adaptive max; Mythos access configurationsource
  • BioMysteryBench Human Difficult44.1%same BioMysteryBench tool setup; difficult subsetsource
  • SpatialBench Verified77.6%life-science evaluation; tool and grader details as specified in the system cardsource
  • ProteinGym Hard49.3%maximum adaptive reasoning; no external tools; protein-design subsetsource
  • Protocol Troubleshooting70.2%bash, file editor, and web search for protocols; maximum adaptive reasoningsource

Gemini 3.8 Flash

Google DeepMind2026-09closed (proprietary)

2 claimed · none independently checked

  • Humanity's Last Exam, HLE-Verified54.9%Google DeepMind product page; reasoning, tools, harness, and date-window not statedsource
  • Maximum output65536 tokensmodel recordsource

Gemini 3.7 Flash

Google DeepMind2026-08closed (proprietary)

9 claimed · none independently checked

  • Artificial Analysis Intelligence Index56 index pointsGoogle comparison table; Aug 2026 snapshot; composition and settings not statedsource
  • FrontierCode 1.1 Main43.6%Google comparison table; production code-quality setting; harness not statedsource
  • DeepSWE v1.165.3%Google comparison table; long-horizon engineering; harness not statedsource
  • Code Arena1588 EloGoogle comparison table; arena protocol not statedsource
  • Terminal-Bench 2.185.8%Google comparison table; agentic terminal coding; harness not statedsource
  • AutomationBench30.4%Google comparison table; private set; harness not statedsource
  • GDPval-AA v21525 EloArtificial Analysis comparison row; task and grader details not stated in the cardsource
  • CharXiv (no tools)84.5%no tools, as named in Google's tablesource
  • CharXiv (with tools)88.7%tools on, as named in Google's tablesource

GPT-5.6 Cyber

OpenAI2026-08closed (proprietary)

0 claimed · none independently checked

    Qwen3.8-Flash-Next

    Qwen2026-08open weights (Qwen Community License 1.0)

    7 claimed · none independently checked

    • DeepSWE v1.158.7%highest across Claude Code and mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
    • SWE-bench Pro62.5%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
    • SWE-bench Multilingual81%mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
    • Toolathlon Verified73.5 % Pass@1Qwen model-card comparison table; default thinking mode; exact harness not statedsource
    • GPQA Diamond91.7%Qwen model-card comparison table; default thinking mode; exact shot count not statedsource
    • HLE35.9%Qwen model-card comparison table; judged by GPT-4o; exact subset and harness not statedsource
    • LiveCodeBench v691.9%Qwen model-card comparison table; exact harness and date window not statedsource

    Qwen3.8-27B

    Qwen2026-08open weights (Apache License 2.0)

    6 claimed · none independently checked

    • Terminal Bench 2.173%Terminus harness; model-card comparison table; other settings not statedsource
    • SWE-bench Pro61.7%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
    • DeepSWE v1.142.2%Claude Code; temperature 1.0, top_p 1.0; 256K contextsource
    • OSWorld-Verified84.3%OSWorld scaffold; model-card table; tool and action settings not statedsource
    • HLE30.8%model-card comparison table; judged by GPT-4o; exact subset not statedsource
    • LiveCodeBench v690.3%model-card comparison table; exact harness and date window not statedsource

    GLM-5.3

    Zhipu AI / Z.ai2026-08open weights (glm-5.3)

    8 claimed · none independently checked

    • Terminal Bench 2.188.2%Claude Code 2.1.207; temperature 1.0, top_p 1.0, max_new_tokens 65,536; 6-hour timeoutsource
    • Terminal Bench 3.028.3%Claude Code 2.1.207; max reasoning; 400K context; 128K output; avg@3; Tool Search disabled; official task verifiersource
    • DeepSWE v1.166.9%mini-SWE-agent; temperature 0.95, top_p 1.0; 6-hour timeout; 400K contextsource
    • FrontierSWE78.1 dominance scoreProximal evaluation; max effort; 1M context; 128K output; score current as of 2026-08-14source
    • SWE-Marathon v1.142.5%Claude Code 2.1.207; max reasoning; temperature 1.0, top_p 0.95; 1M context; 128K outputsource
    • Toolathlon Verified73 % Pass@1official evaluation service; average over 3 independent runssource
    • HLE with tools62.5%temperature 1.0, top_p 0.95; max generation 163,840; max context 300K with context management; GPT-5.6 Luna medium judgesource
    • GDPval-AA v21769 EloArtificial Analysis evaluation; exact task/run details not stated in cardsource

    Gemini 3.6 Flash

    Google DeepMind2026-07closed (proprietary)

    6 claimed · none independently checked

    • SWE-Bench Pro Public58.7%Google comparison table; public split; harness not statedsource
    • DeepSWE v1.149%Google comparison table; harness not statedsource
    • Terminal-Bench 2.178%Terminus-2 harnesssource
    • OSWorld-Verified83%Google comparison table; computer-use setting not otherwise statedsource
    • CharXiv (no tools)85.2%no tools, as named in Google's tablesource
    • CharXiv (with tools)89.4%tools on, as named in Google's tablesource

    Kimi K3

    Moonshot AI2026-07open weights (Kimi K3 License)

    9 claimed · none independently checked

    • GPQA Diamond93.5%max reasoning; temperature 1.0; top_p 0.95 for single-step taskssource
    • HLE-Full43.5%max reasoning; without tools; temperature 1.0; top_p 0.95 for single-step taskssource
    • HLE-Full56%max reasoning; with tools; general tools; temperature 1.0; top_p 0.95 single-step and 1.0 agenticsource
    • DeepSWE v1.167.5%Kimi Code harness; the card says the 67.3 leaderboard result is mini-SWE-agent, but reports 67.5 for K3's runsource
    • Terminal-Bench 2.188.3%Kimi Code harness; max reasoning; agentic top_p 1.0source
    • FrontierSWE81.2 dominance scoreKimi Code harness; dominance recomputed with official script; current as of 2026-07-16source
    • SWE-Marathon42%Claude Code harness; H20-calibrated branch of tasks as of 2026-07-09, before final v1.1source
    • GDPval-AA v21686 EloArtificial Analysis scores cited by Kimi; board observed as of 2026-07-23source
    • Toolathlon-Verified76.5 % Pass@1Kimi comparison table; grader and trial count not stated in the rowsource

    Claude Sonnet 5

    Anthropic2026-06closed (proprietary)

    0 claimed · none independently checked

      MiniMax-M3

      MiniMax2026-06open weights (MiniMax Community License)

      0 claimed · none independently checked

        Claude Opus 4.8

        Anthropic2026-05closed (proprietary)

        0 claimed · none independently checked

          Gemini 3.5 Flash

          Google DeepMind2026-05closed (proprietary)

          5 claimed · none independently checked

          • Terminal-Bench 2.176.2%Terminus-2 harnesssource
          • SWE-Bench Pro Public55.1%single attempt, public splitsource
          • MCP Atlas83.6%Google comparison table; subset/harness not statedsource
          • Toolathlon56.5%Google comparison table; attempt count not statedsource
          • OSWorld-Verified78.4%Google comparison table; computer-use setting not statedsource

          Gemma 4 31B

          Google DeepMind2026-04open weights (Apache License 2.0)

          9 claimed · none independently checked

          • MMLU Pro85.2%instruction-tuned model card result; other harness details not statedsource
          • AIME 202689.2%no toolssource
          • LiveCodeBench v680%instruction-tuned model card result; harness not statedsource
          • Codeforces2150 Elomodel card result; contest protocol not statedsource
          • GPQA Diamond84.3%instruction-tuned model card result; tools not statedsource
          • τ276.9%average over 3source
          • HLE19.5%no toolssource
          • HLE26.5%with searchsource
          • MRCR v2, 8-needle, 128K66.4 % average128K context; 8 needles; average metricsource

          GPT-5.5

          OpenAI2026-04closed (proprietary)

          0 claimed · none independently checked

            GPT-5.5 Pro

            OpenAI2026-04closed (proprietary)

            0 claimed · none independently checked

              DeepSeek-V4-Pro

              DeepSeek2026-04open weights (MIT License)

              8 claimed · none independently checked

              • MMLU90.1 % EMbase model; 5-shotsource
              • MMLU-Pro73.5 % EMbase model; 5-shotsource
              • LongBench-V251.5 % EMbase model; 1-shotsource
              • HLE37.7 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, no tools rowsource
              • HLE48.2 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; with toolssource
              • Terminal Bench 2.067.9 % accuracyDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
              • SWE Verified80.6 % resolvedDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
              • Toolathlon51.8 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource

              DeepSeek-V4-Flash

              DeepSeek2026-04open weights (MIT License)

              0 claimed · none independently checked

                Kimi K2.6

                Moonshot AI2026-04open weights (Modified MIT License)

                6 claimed · none independently checked

                • HLE-Full36.4%max generation 98,304; without toolssource
                • HLE-Full55.5%max generation 98,304; with tools uses search, code interpreter, and web browsing plus context managementsource
                • Terminal-Bench 2.066.7%Terminus-2; preserved thinking; default agent frameworksource
                • SWE-Bench Pro58.6%in-house SWE-agent-derived harness with bash/createfile/insert/view/strreplace/submit toolssource
                • SWE-Bench Verified80.2%same in-house minimal-tool harness; default temperature 1.0, top_p 1.0, 262,144 contextsource
                • LiveCodeBench v689.6%model-card comparison table; exact sampling not stated in rowsource

                MiniMax-M2.7

                MiniMax2026-04research only (NON-COMMERCIAL LICENSE)

                10 claimed · none independently checked

                • MLE Bench Lite66.6 % medal rate22 ML competitions; card describes internal model self-evolution comparisonsource
                • SWE-Pro56.22%MiniMax card; scaffold and run count not stated in the paragraphsource
                • SWE-bench Multilingual76.5%MiniMax card; scaffold and run count not stated in the paragraphsource
                • Multi SWE Bench52.7%MiniMax card; scaffold and run count not stated in the paragraphsource
                • VIBE-Pro55.6%MiniMax card; benchmark harness not stated in the paragraphsource
                • Terminal Bench 257%MiniMax card; benchmark harness not stated in the paragraphsource
                • NL2Repo39.8%MiniMax card; benchmark harness not stated in the paragraphsource
                • GDPval-AA1495 EloMiniMax card; claimed highest among open-weight models; board date not statedsource
                • Toolathlon46.3 % accuracyMiniMax card; benchmark harness and attempt count not statedsource
                • MM Claw62.7%MiniMax card; end-to-end benchmark; harness and attempt count not statedsource

                GPT-5.4 Pro

                OpenAI2026-03closed (proprietary)

                0 claimed · none independently checked

                  GPT-5.4 mini

                  OpenAI2026-03closed (proprietary)

                  0 claimed · none independently checked

                    GPT-5.2-Codex

                    OpenAI2026-01closed (proprietary)

                    0 claimed · none independently checked

                      GPT-5.1-Codex

                      OpenAI2025-11closed (proprietary)

                      0 claimed · none independently checked

                        GPT-5.1-Codex Mini

                        OpenAI2025-11closed (proprietary)

                        0 claimed · none independently checked

                          GPT-5-Codex

                          OpenAI2025-09closed (proprietary)

                          0 claimed · none independently checked

                            o4-mini

                            OpenAI2025-04closed (proprietary)

                            0 claimed · none independently checked

                              02 · The newest

                              The frontier as it stands, Sep 2026

                              The most recently released models, newest first, with every sourced result and the conditions it was produced under. 175 of 175 claims record a harness, effort setting or attempt count — and those conditions routinely move a score further than the model does.The newest models, with what they claim and how the test was run. How you run the test often changes the score more than which model you use.

                              Claude Fable 5.1

                              Anthropic2026-09closed (proprietary)

                              26 claimed · 4 independently checked

                              • SWE-bench Pro81.2 % resolvedadaptive thinking, max effort; average of 5 trialssource
                              • SWE-bench Multilingual89.1 % resolved300 problems across 9 languages; adaptive thinking max; average of 5 trialssource
                              • SWE-bench Multimodal54.7 % resolvedadaptive thinking max; average of 5 trialssource
                              • DeepSWE v1.167.4%113 tasks; adaptive thinking max; average of 5 trials; hidden tests can reject valid but non-reference solutionssource
                              • FrontierCode v1.1 Main50.9%medium reasoning; Anthropic comparison table; older Fable 5 is 53.5 at xhigh, so this is not an apples-to-apples regressionsource
                              • FrontierCode v1.1 Extended63.6%medium reasoning; older Fable 5 is 64.9 at xhigh; the card says some out-of-scope edits are counted as failuressource
                              • FrontierSWE v20.57 dominance scoreProximal agent harness; max reasoning; 5 trials per tasksource
                              • Terminal-Bench 4.055.8%Claude Code with --bare; max thinking; 15 trials per tasksource
                              • Terminal-Bench-Science 0.152.6%Claude Code with --bare; max thinking; 10 trials per task, 700 trials totalsource
                              • CursorBench 3.2.073.4%max effort; 5 runs; Cursor independently reports the benchmark, but no steward check was foundsource
                              • OSWorld 2.077.9 % partial passproduction safeguards active, with Opus 4.8 fallback; default 1080p; max 500 actions; Pass@1, average of 5 runssource
                              • OSWorld 2.041.7 % strict passsame production-safeguard setup as the partial score; results supersede older OSWorld figures and are not comparable to earlier task releasessource
                              • OfficeQA Pro69%Messages API; extracted text sandbox; code execution; output limit 128K; production safeguards and fallbacksource
                              • Legal Agent Benchmark19.09 % all-passpublic Messages API; adaptive max; internal reimplementation; 1,235 problems, 16 defects excluded; n=5source
                              • Legal Agent Benchmark90.81 % criterion passsame internal benchmark and production-safeguard setup; criterion pass is not the all-pass metricsource
                              • GDPval-AA v21853 EloArtificial Analysis; max effort; shell and web; blind pairwise comparison over 220 tasks and 44 occupationssource
                              • Toolathlon Verified77.8 % Pass@1108 tasks; adaptive max; 3 trials each; production safety active with Opus 4.8 fallbacksource
                              • Toolathlon Verified81.5 % Pass@3same 108-task internal harness; three attemptssource
                              • AutomationBench31.4%max effort; private held-out leaderboard; production safeguardssource
                              • ARC-AGI-197.5%max effort; semi-private setsource
                              • ARC-AGI-290%max effort; semi-private setsource
                              • ARC-AGI-197.5%ARC Prize's semi-private evaluation; max effortsource
                              • ARC-AGI-290%ARC Prize's semi-private evaluation; max effortsource
                              • Maximum output128000 tokensmodel recordsource
                              • FrontierMath v2 — Tiers 1-30.9 scoreFable 5.1 max run; verification_code scorersource
                              • FrontierMath v2 — Tier 40.88 scoreFable 5.1 max run; verification_code scorersource

                              Claude Mythos 5.1

                              Anthropic2026-09closed (proprietary)

                              6 claimed · none independently checked

                              • Terminal-Bench 4.060.9%Claude Code with --bare; max thinking; 10 trials per task; Mythos safeguards configurationsource
                              • BioMysteryBench Human Solvable90.3%bash and file editor; restricted domains; adaptive max; Mythos access configurationsource
                              • BioMysteryBench Human Difficult44.1%same BioMysteryBench tool setup; difficult subsetsource
                              • SpatialBench Verified77.6%life-science evaluation; tool and grader details as specified in the system cardsource
                              • ProteinGym Hard49.3%maximum adaptive reasoning; no external tools; protein-design subsetsource
                              • Protocol Troubleshooting70.2%bash, file editor, and web search for protocols; maximum adaptive reasoningsource

                              Gemini 3.8 Flash

                              Google DeepMind2026-09closed (proprietary)

                              2 claimed · none independently checked

                              • Humanity's Last Exam, HLE-Verified54.9%Google DeepMind product page; reasoning, tools, harness, and date-window not statedsource
                              • Maximum output65536 tokensmodel recordsource

                              Gemini 3.7 Flash

                              Google DeepMind2026-08closed (proprietary)

                              9 claimed · none independently checked

                              • Artificial Analysis Intelligence Index56 index pointsGoogle comparison table; Aug 2026 snapshot; composition and settings not statedsource
                              • FrontierCode 1.1 Main43.6%Google comparison table; production code-quality setting; harness not statedsource
                              • DeepSWE v1.165.3%Google comparison table; long-horizon engineering; harness not statedsource
                              • Code Arena1588 EloGoogle comparison table; arena protocol not statedsource
                              • Terminal-Bench 2.185.8%Google comparison table; agentic terminal coding; harness not statedsource
                              • AutomationBench30.4%Google comparison table; private set; harness not statedsource
                              • GDPval-AA v21525 EloArtificial Analysis comparison row; task and grader details not stated in the cardsource
                              • CharXiv (no tools)84.5%no tools, as named in Google's tablesource
                              • CharXiv (with tools)88.7%tools on, as named in Google's tablesource

                              GPT-5.6 Cyber

                              OpenAI2026-08closed (proprietary)

                              0 claimed · none independently checked

                                Qwen3.8-Flash-Next

                                Qwen2026-08open weights (Qwen Community License 1.0)

                                7 claimed · none independently checked

                                • DeepSWE v1.158.7%highest across Claude Code and mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
                                • SWE-bench Pro62.5%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
                                • SWE-bench Multilingual81%mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
                                • Toolathlon Verified73.5 % Pass@1Qwen model-card comparison table; default thinking mode; exact harness not statedsource
                                • GPQA Diamond91.7%Qwen model-card comparison table; default thinking mode; exact shot count not statedsource
                                • HLE35.9%Qwen model-card comparison table; judged by GPT-4o; exact subset and harness not statedsource
                                • LiveCodeBench v691.9%Qwen model-card comparison table; exact harness and date window not statedsource
                                03 · The register

                                Every model we could source, newest first

                                68 models, 175 claims, every one carrying a reachable primary source. Models with no sourced number are absent by design — an announcement is not a result.All 68 models and all 175 numbers, each linked to where it was published. If a company announced a model but published no results, it is not here.

                                Claude Fable 5.1

                                Anthropic2026-09closed (proprietary)

                                26 claimed · 4 independently checked

                                • SWE-bench Pro81.2 % resolvedadaptive thinking, max effort; average of 5 trialssource
                                • SWE-bench Multilingual89.1 % resolved300 problems across 9 languages; adaptive thinking max; average of 5 trialssource
                                • SWE-bench Multimodal54.7 % resolvedadaptive thinking max; average of 5 trialssource
                                • DeepSWE v1.167.4%113 tasks; adaptive thinking max; average of 5 trials; hidden tests can reject valid but non-reference solutionssource
                                • FrontierCode v1.1 Main50.9%medium reasoning; Anthropic comparison table; older Fable 5 is 53.5 at xhigh, so this is not an apples-to-apples regressionsource
                                • FrontierCode v1.1 Extended63.6%medium reasoning; older Fable 5 is 64.9 at xhigh; the card says some out-of-scope edits are counted as failuressource
                                • FrontierSWE v20.57 dominance scoreProximal agent harness; max reasoning; 5 trials per tasksource
                                • Terminal-Bench 4.055.8%Claude Code with --bare; max thinking; 15 trials per tasksource
                                • Terminal-Bench-Science 0.152.6%Claude Code with --bare; max thinking; 10 trials per task, 700 trials totalsource
                                • CursorBench 3.2.073.4%max effort; 5 runs; Cursor independently reports the benchmark, but no steward check was foundsource
                                • OSWorld 2.077.9 % partial passproduction safeguards active, with Opus 4.8 fallback; default 1080p; max 500 actions; Pass@1, average of 5 runssource
                                • OSWorld 2.041.7 % strict passsame production-safeguard setup as the partial score; results supersede older OSWorld figures and are not comparable to earlier task releasessource
                                • OfficeQA Pro69%Messages API; extracted text sandbox; code execution; output limit 128K; production safeguards and fallbacksource
                                • Legal Agent Benchmark19.09 % all-passpublic Messages API; adaptive max; internal reimplementation; 1,235 problems, 16 defects excluded; n=5source
                                • Legal Agent Benchmark90.81 % criterion passsame internal benchmark and production-safeguard setup; criterion pass is not the all-pass metricsource
                                • GDPval-AA v21853 EloArtificial Analysis; max effort; shell and web; blind pairwise comparison over 220 tasks and 44 occupationssource
                                • Toolathlon Verified77.8 % Pass@1108 tasks; adaptive max; 3 trials each; production safety active with Opus 4.8 fallbacksource
                                • Toolathlon Verified81.5 % Pass@3same 108-task internal harness; three attemptssource
                                • AutomationBench31.4%max effort; private held-out leaderboard; production safeguardssource
                                • ARC-AGI-197.5%max effort; semi-private setsource
                                • ARC-AGI-290%max effort; semi-private setsource
                                • ARC-AGI-197.5%ARC Prize's semi-private evaluation; max effortsource
                                • ARC-AGI-290%ARC Prize's semi-private evaluation; max effortsource
                                • Maximum output128000 tokensmodel recordsource
                                • FrontierMath v2 — Tiers 1-30.9 scoreFable 5.1 max run; verification_code scorersource
                                • FrontierMath v2 — Tier 40.88 scoreFable 5.1 max run; verification_code scorersource

                                Claude Mythos 5.1

                                Anthropic2026-09closed (proprietary)

                                6 claimed · none independently checked

                                • Terminal-Bench 4.060.9%Claude Code with --bare; max thinking; 10 trials per task; Mythos safeguards configurationsource
                                • BioMysteryBench Human Solvable90.3%bash and file editor; restricted domains; adaptive max; Mythos access configurationsource
                                • BioMysteryBench Human Difficult44.1%same BioMysteryBench tool setup; difficult subsetsource
                                • SpatialBench Verified77.6%life-science evaluation; tool and grader details as specified in the system cardsource
                                • ProteinGym Hard49.3%maximum adaptive reasoning; no external tools; protein-design subsetsource
                                • Protocol Troubleshooting70.2%bash, file editor, and web search for protocols; maximum adaptive reasoningsource

                                Gemini 3.8 Flash

                                Google DeepMind2026-09closed (proprietary)

                                2 claimed · none independently checked

                                • Humanity's Last Exam, HLE-Verified54.9%Google DeepMind product page; reasoning, tools, harness, and date-window not statedsource
                                • Maximum output65536 tokensmodel recordsource

                                Gemini 3.7 Flash

                                Google DeepMind2026-08closed (proprietary)

                                9 claimed · none independently checked

                                • Artificial Analysis Intelligence Index56 index pointsGoogle comparison table; Aug 2026 snapshot; composition and settings not statedsource
                                • FrontierCode 1.1 Main43.6%Google comparison table; production code-quality setting; harness not statedsource
                                • DeepSWE v1.165.3%Google comparison table; long-horizon engineering; harness not statedsource
                                • Code Arena1588 EloGoogle comparison table; arena protocol not statedsource
                                • Terminal-Bench 2.185.8%Google comparison table; agentic terminal coding; harness not statedsource
                                • AutomationBench30.4%Google comparison table; private set; harness not statedsource
                                • GDPval-AA v21525 EloArtificial Analysis comparison row; task and grader details not stated in the cardsource
                                • CharXiv (no tools)84.5%no tools, as named in Google's tablesource
                                • CharXiv (with tools)88.7%tools on, as named in Google's tablesource

                                GPT-5.6 Cyber

                                OpenAI2026-08closed (proprietary)

                                0 claimed · none independently checked

                                  Qwen3.8-Flash-Next

                                  Qwen2026-08open weights (Qwen Community License 1.0)

                                  7 claimed · none independently checked

                                  • DeepSWE v1.158.7%highest across Claude Code and mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
                                  • SWE-bench Pro62.5%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
                                  • SWE-bench Multilingual81%mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
                                  • Toolathlon Verified73.5 % Pass@1Qwen model-card comparison table; default thinking mode; exact harness not statedsource
                                  • GPQA Diamond91.7%Qwen model-card comparison table; default thinking mode; exact shot count not statedsource
                                  • HLE35.9%Qwen model-card comparison table; judged by GPT-4o; exact subset and harness not statedsource
                                  • LiveCodeBench v691.9%Qwen model-card comparison table; exact harness and date window not statedsource

                                  Qwen3.8-27B

                                  Qwen2026-08open weights (Apache License 2.0)

                                  6 claimed · none independently checked

                                  • Terminal Bench 2.173%Terminus harness; model-card comparison table; other settings not statedsource
                                  • SWE-bench Pro61.7%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
                                  • DeepSWE v1.142.2%Claude Code; temperature 1.0, top_p 1.0; 256K contextsource
                                  • OSWorld-Verified84.3%OSWorld scaffold; model-card table; tool and action settings not statedsource
                                  • HLE30.8%model-card comparison table; judged by GPT-4o; exact subset not statedsource
                                  • LiveCodeBench v690.3%model-card comparison table; exact harness and date window not statedsource

                                  GLM-5.3

                                  Zhipu AI / Z.ai2026-08open weights (glm-5.3)

                                  8 claimed · none independently checked

                                  • Terminal Bench 2.188.2%Claude Code 2.1.207; temperature 1.0, top_p 1.0, max_new_tokens 65,536; 6-hour timeoutsource
                                  • Terminal Bench 3.028.3%Claude Code 2.1.207; max reasoning; 400K context; 128K output; avg@3; Tool Search disabled; official task verifiersource
                                  • DeepSWE v1.166.9%mini-SWE-agent; temperature 0.95, top_p 1.0; 6-hour timeout; 400K contextsource
                                  • FrontierSWE78.1 dominance scoreProximal evaluation; max effort; 1M context; 128K output; score current as of 2026-08-14source
                                  • SWE-Marathon v1.142.5%Claude Code 2.1.207; max reasoning; temperature 1.0, top_p 0.95; 1M context; 128K outputsource
                                  • Toolathlon Verified73 % Pass@1official evaluation service; average over 3 independent runssource
                                  • HLE with tools62.5%temperature 1.0, top_p 0.95; max generation 163,840; max context 300K with context management; GPT-5.6 Luna medium judgesource
                                  • GDPval-AA v21769 EloArtificial Analysis evaluation; exact task/run details not stated in cardsource

                                  Grok 4.5

                                  xAI2026-08closed

                                  1 claimed · 1 independently checked

                                  • ARC-AGI-252.6%ARC Prize independent evaluationsource

                                  Grok 4.6

                                  xAI2026-08closed

                                  1 claimed · 1 independently checked

                                  • ARC-AGI-267.1%ARC Prize independent evaluationsource

                                  Claude Opus 5

                                  Anthropic2026-07closed (proprietary)

                                  3 claimed · 3 independently checked

                                  • ARC-AGI-197.5%max effortsource
                                  • ARC-AGI-290.4%max effortsource
                                  • ARC-AGI-330.16%high effortsource

                                  Gemini 3.6 Flash

                                  Google DeepMind2026-07closed (proprietary)

                                  6 claimed · none independently checked

                                  • SWE-Bench Pro Public58.7%Google comparison table; public split; harness not statedsource
                                  • DeepSWE v1.149%Google comparison table; harness not statedsource
                                  • Terminal-Bench 2.178%Terminus-2 harnesssource
                                  • OSWorld-Verified83%Google comparison table; computer-use setting not otherwise statedsource
                                  • CharXiv (no tools)85.2%no tools, as named in Google's tablesource
                                  • CharXiv (with tools)89.4%tools on, as named in Google's tablesource

                                  GPT-5.6 Sol

                                  OpenAI2026-07closed (proprietary)

                                  3 claimed · 3 independently checked

                                  • ARC-AGI-196.5%max reasoning effort; semi-privatesource
                                  • ARC-AGI-292.5%max reasoning effort; semi-privatesource
                                  • ARC-AGI-37.78%max reasoning effort; semi-privatesource

                                  GPT-5.6 Terra

                                  OpenAI2026-07closed (proprietary)

                                  3 claimed · 3 independently checked

                                  • ARC-AGI-196.5%max reasoning effort; semi-privatesource
                                  • ARC-AGI-283.9%max reasoning effort; semi-privatesource
                                  • ARC-AGI-30.8%max reasoning effort; semi-privatesource

                                  GPT-5.6 Luna

                                  OpenAI2026-07closed (proprietary)

                                  3 claimed · 3 independently checked

                                  • ARC-AGI-188%max reasoning effort; semi-privatesource
                                  • ARC-AGI-259.5%max reasoning effort; semi-privatesource
                                  • ARC-AGI-30.18%max reasoning effort; semi-privatesource

                                  Kimi K3

                                  Moonshot AI2026-07open weights (Kimi K3 License)

                                  9 claimed · none independently checked

                                  • GPQA Diamond93.5%max reasoning; temperature 1.0; top_p 0.95 for single-step taskssource
                                  • HLE-Full43.5%max reasoning; without tools; temperature 1.0; top_p 0.95 for single-step taskssource
                                  • HLE-Full56%max reasoning; with tools; general tools; temperature 1.0; top_p 0.95 single-step and 1.0 agenticsource
                                  • DeepSWE v1.167.5%Kimi Code harness; the card says the 67.3 leaderboard result is mini-SWE-agent, but reports 67.5 for K3's runsource
                                  • Terminal-Bench 2.188.3%Kimi Code harness; max reasoning; agentic top_p 1.0source
                                  • FrontierSWE81.2 dominance scoreKimi Code harness; dominance recomputed with official script; current as of 2026-07-16source
                                  • SWE-Marathon42%Claude Code harness; H20-calibrated branch of tasks as of 2026-07-09, before final v1.1source
                                  • GDPval-AA v21686 EloArtificial Analysis scores cited by Kimi; board observed as of 2026-07-23source
                                  • Toolathlon-Verified76.5 % Pass@1Kimi comparison table; grader and trial count not stated in the rowsource

                                  Claude Sonnet 5

                                  Anthropic2026-06closed (proprietary)

                                  0 claimed · none independently checked

                                    MiniMax-M3

                                    MiniMax2026-06open weights (MiniMax Community License)

                                    0 claimed · none independently checked

                                      Claude Opus 4.8

                                      Anthropic2026-05closed (proprietary)

                                      0 claimed · none independently checked

                                        Gemini 3.5 Flash

                                        Google DeepMind2026-05closed (proprietary)

                                        5 claimed · none independently checked

                                        • Terminal-Bench 2.176.2%Terminus-2 harnesssource
                                        • SWE-Bench Pro Public55.1%single attempt, public splitsource
                                        • MCP Atlas83.6%Google comparison table; subset/harness not statedsource
                                        • Toolathlon56.5%Google comparison table; attempt count not statedsource
                                        • OSWorld-Verified78.4%Google comparison table; computer-use setting not statedsource

                                        Gemma 4 31B

                                        Google DeepMind2026-04open weights (Apache License 2.0)

                                        9 claimed · none independently checked

                                        • MMLU Pro85.2%instruction-tuned model card result; other harness details not statedsource
                                        • AIME 202689.2%no toolssource
                                        • LiveCodeBench v680%instruction-tuned model card result; harness not statedsource
                                        • Codeforces2150 Elomodel card result; contest protocol not statedsource
                                        • GPQA Diamond84.3%instruction-tuned model card result; tools not statedsource
                                        • τ276.9%average over 3source
                                        • HLE19.5%no toolssource
                                        • HLE26.5%with searchsource
                                        • MRCR v2, 8-needle, 128K66.4 % average128K context; 8 needles; average metricsource

                                        GPT-5.5

                                        OpenAI2026-04closed (proprietary)

                                        0 claimed · none independently checked

                                          GPT-5.5 Pro

                                          OpenAI2026-04closed (proprietary)

                                          0 claimed · none independently checked

                                            DeepSeek-V4-Pro

                                            DeepSeek2026-04open weights (MIT License)

                                            8 claimed · none independently checked

                                            • MMLU90.1 % EMbase model; 5-shotsource
                                            • MMLU-Pro73.5 % EMbase model; 5-shotsource
                                            • LongBench-V251.5 % EMbase model; 1-shotsource
                                            • HLE37.7 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, no tools rowsource
                                            • HLE48.2 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; with toolssource
                                            • Terminal Bench 2.067.9 % accuracyDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
                                            • SWE Verified80.6 % resolvedDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
                                            • Toolathlon51.8 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource

                                            DeepSeek-V4-Flash

                                            DeepSeek2026-04open weights (MIT License)

                                            0 claimed · none independently checked

                                              Kimi K2.6

                                              Moonshot AI2026-04open weights (Modified MIT License)

                                              6 claimed · none independently checked

                                              • HLE-Full36.4%max generation 98,304; without toolssource
                                              • HLE-Full55.5%max generation 98,304; with tools uses search, code interpreter, and web browsing plus context managementsource
                                              • Terminal-Bench 2.066.7%Terminus-2; preserved thinking; default agent frameworksource
                                              • SWE-Bench Pro58.6%in-house SWE-agent-derived harness with bash/createfile/insert/view/strreplace/submit toolssource
                                              • SWE-Bench Verified80.2%same in-house minimal-tool harness; default temperature 1.0, top_p 1.0, 262,144 contextsource
                                              • LiveCodeBench v689.6%model-card comparison table; exact sampling not stated in rowsource

                                              MiniMax-M2.7

                                              MiniMax2026-04research only (NON-COMMERCIAL LICENSE)

                                              10 claimed · none independently checked

                                              • MLE Bench Lite66.6 % medal rate22 ML competitions; card describes internal model self-evolution comparisonsource
                                              • SWE-Pro56.22%MiniMax card; scaffold and run count not stated in the paragraphsource
                                              • SWE-bench Multilingual76.5%MiniMax card; scaffold and run count not stated in the paragraphsource
                                              • Multi SWE Bench52.7%MiniMax card; scaffold and run count not stated in the paragraphsource
                                              • VIBE-Pro55.6%MiniMax card; benchmark harness not stated in the paragraphsource
                                              • Terminal Bench 257%MiniMax card; benchmark harness not stated in the paragraphsource
                                              • NL2Repo39.8%MiniMax card; benchmark harness not stated in the paragraphsource
                                              • GDPval-AA1495 EloMiniMax card; claimed highest among open-weight models; board date not statedsource
                                              • Toolathlon46.3 % accuracyMiniMax card; benchmark harness and attempt count not statedsource
                                              • MM Claw62.7%MiniMax card; end-to-end benchmark; harness and attempt count not statedsource

                                              Claude Mythos Preview (early)

                                              Anthropic2026-04

                                              1 claimed · 1 independently checked

                                              • METR time horizon (TH1.1)1044.78 minutesp50 horizon length; METR-Horizon-v1.1source

                                              GPT-5.4

                                              OpenAI2026-03closed (proprietary)

                                              1 claimed · 1 independently checked

                                              • METR time horizon (TH1.1)341.74 minutesp50 horizon length; METR-Horizon-v1.1source

                                              GPT-5.4 Pro

                                              OpenAI2026-03closed (proprietary)

                                              0 claimed · none independently checked

                                                GPT-5.4 mini

                                                OpenAI2026-03closed (proprietary)

                                                0 claimed · none independently checked

                                                  Gemini 3.1 Pro

                                                  Google DeepMind2026-02closed (proprietary)

                                                  6 claimed · 1 independently checked

                                                  • HLE (no tools)44.4%no toolssource
                                                  • HLE (Search + Code)51.4%Search blocklist + Codesource
                                                  • GPQA Diamond (no tools)94.3%no toolssource
                                                  • Terminal-Bench 2.068.5%Terminus-2 harnesssource
                                                  • SWE-Bench Verified80.6%single attemptsource
                                                  • METR time horizon (TH1.1)384.15 minutesp50 horizon length; METR-Horizon-v1.1source

                                                  GPT-5.3-Codex

                                                  OpenAI2026-02closed (proprietary)

                                                  1 claimed · 1 independently checked

                                                  • METR time horizon (TH1.1)349.53 minutesp50 horizon length; METR-Horizon-v1.1source

                                                  Claude Opus 4.6

                                                  Anthropic2026-02

                                                  1 claimed · 1 independently checked

                                                  • METR time horizon (TH1.1)718.81 minutesp50 horizon length; METR-Horizon-v1.1source

                                                  GPT-5.2-Codex

                                                  OpenAI2026-01closed (proprietary)

                                                  0 claimed · none independently checked

                                                    GPT-5.1-Codex-Max

                                                    OpenAI2025-12closed (proprietary)

                                                    1 claimed · 1 independently checked

                                                    • METR time horizon (TH1.1)223.71 minutesp50 horizon length; METR-Horizon-v1.1source

                                                    GPT-5.2

                                                    OpenAI2025-12

                                                    1 claimed · 1 independently checked

                                                    • METR time horizon (TH1.1)352.25 minutesp50 horizon length; METR-Horizon-v1.1source

                                                    Devstral 2 123B

                                                    Mistral2025-12

                                                    2 claimed · 1 independently checked

                                                    • SWE-bench Verified72.2%Mistral-reported resultsource
                                                    • SWE-bench Verified53.8%mini-SWE-agent harness; checked: truesource

                                                    Devstral Small 2 24B

                                                    Mistral2025-12open weights (Apache 2.0)

                                                    2 claimed · 1 independently checked

                                                    • SWE-bench Verified68%Mistral-reported resultsource
                                                    • SWE-bench Verified56.4%mini-SWE-agent harness; checked: truesource

                                                    GPT-5.1-Codex

                                                    OpenAI2025-11closed (proprietary)

                                                    0 claimed · none independently checked

                                                      GPT-5.1-Codex Mini

                                                      OpenAI2025-11closed (proprietary)

                                                      0 claimed · none independently checked

                                                        Gemini 3 Pro

                                                        Google DeepMind2025-11

                                                        1 claimed · 1 independently checked

                                                        • METR time horizon (TH1.1)224.33 minutesp50 horizon length; METR-Horizon-v1.1source

                                                        Claude Opus 4.5

                                                        Anthropic2025-11

                                                        1 claimed · 1 independently checked

                                                        • METR time horizon (TH1.1)292.99 minutesp50 horizon length; METR-Horizon-v1.1source

                                                        GPT-5-Codex

                                                        OpenAI2025-09closed (proprietary)

                                                        0 claimed · none independently checked

                                                          Magistral Small 1.2

                                                          Mistral2025-09open weights (Apache 2.0)

                                                          2 claimed · 1 independently checked

                                                          • GPQA Diamond70.07%Mistral-reported resultsource
                                                          • GPQA Diamond47.6%Epoch AI independent runsource

                                                          Claude Opus 4.1

                                                          Anthropic2025-08

                                                          1 claimed · 1 independently checked

                                                          • METR time horizon (TH1.1)100.47 minutesp50 horizon length; METR-Horizon-v1.1source

                                                          GPT-5

                                                          OpenAI2025-08

                                                          1 claimed · 1 independently checked

                                                          • METR time horizon (TH1.1)203.01 minutesp50 horizon length; METR-Horizon-v1.1source

                                                          Devstral Small 1.1

                                                          Mistral2025-07open weights (Apache 2.0)

                                                          2 claimed · 2 independently checked

                                                          • SWE-bench Verified53.6%OpenHands harness; checked: truesource
                                                          • SWE-bench Verified38%SWE-agent harness; checked: truesource

                                                          Claude Opus 4

                                                          Anthropic2025-05

                                                          1 claimed · 1 independently checked

                                                          • METR time horizon (TH1.1)100.37 minutesp50 horizon length; METR-Horizon-v1.1source

                                                          o3

                                                          OpenAI2025-04closed (proprietary)

                                                          1 claimed · 1 independently checked

                                                          • METR time horizon (TH1.1)119.73 minutesp50 horizon length; METR-Horizon-v1.1source

                                                          o4-mini

                                                          OpenAI2025-04closed (proprietary)

                                                          0 claimed · none independently checked

                                                            Llama 4 Maverick

                                                            Meta2025-04

                                                            1 claimed · 1 independently checked

                                                            • SWE-bench Verified21.04%checked: truesource

                                                            Llama 4 Scout

                                                            Meta2025-04

                                                            1 claimed · 1 independently checked

                                                            • SWE-bench Verified9.06%checked: truesource

                                                            Mistral Small 3.1

                                                            Mistral2025-03open weights (Apache 2.0)

                                                            2 claimed · 1 independently checked

                                                            • GPQA Diamond45.96%Mistral-reported resultsource
                                                            • GPQA Diamond41.9%Epoch AI independent runsource

                                                            Claude 3.7 Sonnet

                                                            Anthropic2025-02

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)60.39 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            Grok 3

                                                            xAI2025-02closed

                                                            1 claimed · 1 independently checked

                                                            • ARC-AGI-20%ARC Prize independent evaluationsource

                                                            o1

                                                            OpenAI2024-12

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)38.83 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            Claude 3.5 Sonnet (Oct 2024)

                                                            Anthropic2024-10

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)20.52 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            o1-preview

                                                            OpenAI2024-09

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)20.33 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            Claude 3.5 Sonnet (Jun 2024)

                                                            Anthropic2024-06

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)11.4 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-4o

                                                            OpenAI2024-05

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)6.99 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-4 Turbo

                                                            OpenAI2024-04

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)3.73 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            Claude 3 Opus

                                                            Anthropic2024-03

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)3.95 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-4 (1106)

                                                            OpenAI2023-11

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)4.04 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-4

                                                            OpenAI2023-03

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)3.99 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-3.5 Turbo Instruct

                                                            OpenAI2022-03

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)0.6 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-3 (davinci-002)

                                                            OpenAI2020-05

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)0.14 minutesp50 horizon length; METR-Horizon-v1.1source

                                                            GPT-2

                                                            OpenAI2019-02

                                                            1 claimed · 1 independently checked

                                                            • METR time horizon (TH1.1)0.05 minutesp50 horizon length; METR-Horizon-v1.1source
                                                            04 · How this was built

                                                            What counts as checked, and what we left out

                                                            Every number was read from a primary source — a lab's own model card, system card or pricing page, or a steward's published results. Nothing came from an aggregator or a news write-up. A result counts as checkedonly when a benchmark's own steward published or re-ran it, and that steward is always named. Numbers a lab published only as an image or a chart are recorded as unreadable rather than guessed at, and models with no sourced figure are omitted entirely. Observed 2026-09-02; the full working, including everything we could not verify, is in the repository.