175 claims about 68 models. 53 have been checked by anyone else.
Every published benchmark result we could source for the current frontier and the open-weight field, tagged by who stands behind it. 30% carry an independent steward's check; the rest rest on the word of the party with an interest in the number. Neither is dismissed here — but they are never drawn the same way.What AI companies say their newest models can do — and whether anyone independent has actually tested it. Usually, nobody has. We show both, and we never draw them the same way.
Solid = an independent scorekeeper checked it. Hollow= the lab's own claim.
27 of 68 models have no independent result at all
Independent checks come from a short list of stewards: ARC Prize Foundation, Epoch AI, METR, SWE-bench maintainers. Where a model is absent from all of them, every number it has is its own. A leaderboard row that is merely submittedis not a check — SWE-bench's current entries sit unexamined, and are counted here as claims.Only a handful of independent groups actually re-run these tests: ARC Prize Foundation, Epoch AI, METR, SWE-bench maintainers. If a model is missing from all of them, every number about it came from the company that built it.
Claude Mythos 5.1
Anthropic2026-09closed (proprietary)
6 claimed · none independently checked
- Terminal-Bench 4.060.9%Claude Code with --bare; max thinking; 10 trials per task; Mythos safeguards configurationsource
- BioMysteryBench Human Solvable90.3%bash and file editor; restricted domains; adaptive max; Mythos access configurationsource
- BioMysteryBench Human Difficult44.1%same BioMysteryBench tool setup; difficult subsetsource
- SpatialBench Verified77.6%life-science evaluation; tool and grader details as specified in the system cardsource
- ProteinGym Hard49.3%maximum adaptive reasoning; no external tools; protein-design subsetsource
- Protocol Troubleshooting70.2%bash, file editor, and web search for protocols; maximum adaptive reasoningsource
Gemini 3.8 Flash
Google DeepMind2026-09closed (proprietary)
2 claimed · none independently checked
Gemini 3.7 Flash
Google DeepMind2026-08closed (proprietary)
9 claimed · none independently checked
- Artificial Analysis Intelligence Index56 index pointsGoogle comparison table; Aug 2026 snapshot; composition and settings not statedsource
- FrontierCode 1.1 Main43.6%Google comparison table; production code-quality setting; harness not statedsource
- DeepSWE v1.165.3%Google comparison table; long-horizon engineering; harness not statedsource
- Code Arena1588 EloGoogle comparison table; arena protocol not statedsource
- Terminal-Bench 2.185.8%Google comparison table; agentic terminal coding; harness not statedsource
- AutomationBench30.4%Google comparison table; private set; harness not statedsource
- GDPval-AA v21525 EloArtificial Analysis comparison row; task and grader details not stated in the cardsource
- CharXiv (no tools)84.5%no tools, as named in Google's tablesource
- CharXiv (with tools)88.7%tools on, as named in Google's tablesource
GPT-5.6 Cyber
OpenAI2026-08closed (proprietary)
0 claimed · none independently checked
Qwen3.8-Flash-Next
Qwen2026-08open weights (Qwen Community License 1.0)
7 claimed · none independently checked
- DeepSWE v1.158.7%highest across Claude Code and mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
- SWE-bench Pro62.5%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
- SWE-bench Multilingual81%mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
- Toolathlon Verified73.5 % Pass@1Qwen model-card comparison table; default thinking mode; exact harness not statedsource
- GPQA Diamond91.7%Qwen model-card comparison table; default thinking mode; exact shot count not statedsource
- HLE35.9%Qwen model-card comparison table; judged by GPT-4o; exact subset and harness not statedsource
- LiveCodeBench v691.9%Qwen model-card comparison table; exact harness and date window not statedsource
Qwen3.8-27B
Qwen2026-08open weights (Apache License 2.0)
6 claimed · none independently checked
- Terminal Bench 2.173%Terminus harness; model-card comparison table; other settings not statedsource
- SWE-bench Pro61.7%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
- DeepSWE v1.142.2%Claude Code; temperature 1.0, top_p 1.0; 256K contextsource
- OSWorld-Verified84.3%OSWorld scaffold; model-card table; tool and action settings not statedsource
- HLE30.8%model-card comparison table; judged by GPT-4o; exact subset not statedsource
- LiveCodeBench v690.3%model-card comparison table; exact harness and date window not statedsource
GLM-5.3
Zhipu AI / Z.ai2026-08open weights (glm-5.3)
8 claimed · none independently checked
- Terminal Bench 2.188.2%Claude Code 2.1.207; temperature 1.0, top_p 1.0, max_new_tokens 65,536; 6-hour timeoutsource
- Terminal Bench 3.028.3%Claude Code 2.1.207; max reasoning; 400K context; 128K output; avg@3; Tool Search disabled; official task verifiersource
- DeepSWE v1.166.9%mini-SWE-agent; temperature 0.95, top_p 1.0; 6-hour timeout; 400K contextsource
- FrontierSWE78.1 dominance scoreProximal evaluation; max effort; 1M context; 128K output; score current as of 2026-08-14source
- SWE-Marathon v1.142.5%Claude Code 2.1.207; max reasoning; temperature 1.0, top_p 0.95; 1M context; 128K outputsource
- Toolathlon Verified73 % Pass@1official evaluation service; average over 3 independent runssource
- HLE with tools62.5%temperature 1.0, top_p 0.95; max generation 163,840; max context 300K with context management; GPT-5.6 Luna medium judgesource
- GDPval-AA v21769 EloArtificial Analysis evaluation; exact task/run details not stated in cardsource
Gemini 3.6 Flash
Google DeepMind2026-07closed (proprietary)
6 claimed · none independently checked
- SWE-Bench Pro Public58.7%Google comparison table; public split; harness not statedsource
- DeepSWE v1.149%Google comparison table; harness not statedsource
- Terminal-Bench 2.178%Terminus-2 harnesssource
- OSWorld-Verified83%Google comparison table; computer-use setting not otherwise statedsource
- CharXiv (no tools)85.2%no tools, as named in Google's tablesource
- CharXiv (with tools)89.4%tools on, as named in Google's tablesource
Kimi K3
Moonshot AI2026-07open weights (Kimi K3 License)
9 claimed · none independently checked
- GPQA Diamond93.5%max reasoning; temperature 1.0; top_p 0.95 for single-step taskssource
- HLE-Full43.5%max reasoning; without tools; temperature 1.0; top_p 0.95 for single-step taskssource
- HLE-Full56%max reasoning; with tools; general tools; temperature 1.0; top_p 0.95 single-step and 1.0 agenticsource
- DeepSWE v1.167.5%Kimi Code harness; the card says the 67.3 leaderboard result is mini-SWE-agent, but reports 67.5 for K3's runsource
- Terminal-Bench 2.188.3%Kimi Code harness; max reasoning; agentic top_p 1.0source
- FrontierSWE81.2 dominance scoreKimi Code harness; dominance recomputed with official script; current as of 2026-07-16source
- SWE-Marathon42%Claude Code harness; H20-calibrated branch of tasks as of 2026-07-09, before final v1.1source
- GDPval-AA v21686 EloArtificial Analysis scores cited by Kimi; board observed as of 2026-07-23source
- Toolathlon-Verified76.5 % Pass@1Kimi comparison table; grader and trial count not stated in the rowsource
Claude Sonnet 5
Anthropic2026-06closed (proprietary)
0 claimed · none independently checked
MiniMax-M3
MiniMax2026-06open weights (MiniMax Community License)
0 claimed · none independently checked
Claude Opus 4.8
Anthropic2026-05closed (proprietary)
0 claimed · none independently checked
Gemini 3.5 Flash
Google DeepMind2026-05closed (proprietary)
5 claimed · none independently checked
- Terminal-Bench 2.176.2%Terminus-2 harnesssource
- SWE-Bench Pro Public55.1%single attempt, public splitsource
- MCP Atlas83.6%Google comparison table; subset/harness not statedsource
- Toolathlon56.5%Google comparison table; attempt count not statedsource
- OSWorld-Verified78.4%Google comparison table; computer-use setting not statedsource
Gemma 4 31B
Google DeepMind2026-04open weights (Apache License 2.0)
9 claimed · none independently checked
- MMLU Pro85.2%instruction-tuned model card result; other harness details not statedsource
- AIME 202689.2%no toolssource
- LiveCodeBench v680%instruction-tuned model card result; harness not statedsource
- Codeforces2150 Elomodel card result; contest protocol not statedsource
- GPQA Diamond84.3%instruction-tuned model card result; tools not statedsource
- τ276.9%average over 3source
- HLE19.5%no toolssource
- HLE26.5%with searchsource
- MRCR v2, 8-needle, 128K66.4 % average128K context; 8 needles; average metricsource
GPT-5.5
OpenAI2026-04closed (proprietary)
0 claimed · none independently checked
GPT-5.5 Pro
OpenAI2026-04closed (proprietary)
0 claimed · none independently checked
DeepSeek-V4-Pro
DeepSeek2026-04open weights (MIT License)
8 claimed · none independently checked
- MMLU90.1 % EMbase model; 5-shotsource
- MMLU-Pro73.5 % EMbase model; 5-shotsource
- LongBench-V251.5 % EMbase model; 1-shotsource
- HLE37.7 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, no tools rowsource
- HLE48.2 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; with toolssource
- Terminal Bench 2.067.9 % accuracyDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
- SWE Verified80.6 % resolvedDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
- Toolathlon51.8 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
DeepSeek-V4-Flash
DeepSeek2026-04open weights (MIT License)
0 claimed · none independently checked
Kimi K2.6
Moonshot AI2026-04open weights (Modified MIT License)
6 claimed · none independently checked
- HLE-Full36.4%max generation 98,304; without toolssource
- HLE-Full55.5%max generation 98,304; with tools uses search, code interpreter, and web browsing plus context managementsource
- Terminal-Bench 2.066.7%Terminus-2; preserved thinking; default agent frameworksource
- SWE-Bench Pro58.6%in-house SWE-agent-derived harness with bash/createfile/insert/view/strreplace/submit toolssource
- SWE-Bench Verified80.2%same in-house minimal-tool harness; default temperature 1.0, top_p 1.0, 262,144 contextsource
- LiveCodeBench v689.6%model-card comparison table; exact sampling not stated in rowsource
MiniMax-M2.7
MiniMax2026-04research only (NON-COMMERCIAL LICENSE)
10 claimed · none independently checked
- MLE Bench Lite66.6 % medal rate22 ML competitions; card describes internal model self-evolution comparisonsource
- SWE-Pro56.22%MiniMax card; scaffold and run count not stated in the paragraphsource
- SWE-bench Multilingual76.5%MiniMax card; scaffold and run count not stated in the paragraphsource
- Multi SWE Bench52.7%MiniMax card; scaffold and run count not stated in the paragraphsource
- VIBE-Pro55.6%MiniMax card; benchmark harness not stated in the paragraphsource
- Terminal Bench 257%MiniMax card; benchmark harness not stated in the paragraphsource
- NL2Repo39.8%MiniMax card; benchmark harness not stated in the paragraphsource
- GDPval-AA1495 EloMiniMax card; claimed highest among open-weight models; board date not statedsource
- Toolathlon46.3 % accuracyMiniMax card; benchmark harness and attempt count not statedsource
- MM Claw62.7%MiniMax card; end-to-end benchmark; harness and attempt count not statedsource
GPT-5.4 Pro
OpenAI2026-03closed (proprietary)
0 claimed · none independently checked
GPT-5.4 mini
OpenAI2026-03closed (proprietary)
0 claimed · none independently checked
GPT-5.2-Codex
OpenAI2026-01closed (proprietary)
0 claimed · none independently checked
GPT-5.1-Codex
OpenAI2025-11closed (proprietary)
0 claimed · none independently checked
GPT-5.1-Codex Mini
OpenAI2025-11closed (proprietary)
0 claimed · none independently checked
GPT-5-Codex
OpenAI2025-09closed (proprietary)
0 claimed · none independently checked
o4-mini
OpenAI2025-04closed (proprietary)
0 claimed · none independently checked
The frontier as it stands, Sep 2026
The most recently released models, newest first, with every sourced result and the conditions it was produced under. 175 of 175 claims record a harness, effort setting or attempt count — and those conditions routinely move a score further than the model does.The newest models, with what they claim and how the test was run. How you run the test often changes the score more than which model you use.
Claude Fable 5.1
Anthropic2026-09closed (proprietary)
26 claimed · 4 independently checked
- SWE-bench Pro81.2 % resolvedadaptive thinking, max effort; average of 5 trialssource
- SWE-bench Multilingual89.1 % resolved300 problems across 9 languages; adaptive thinking max; average of 5 trialssource
- SWE-bench Multimodal54.7 % resolvedadaptive thinking max; average of 5 trialssource
- DeepSWE v1.167.4%113 tasks; adaptive thinking max; average of 5 trials; hidden tests can reject valid but non-reference solutionssource
- FrontierCode v1.1 Main50.9%medium reasoning; Anthropic comparison table; older Fable 5 is 53.5 at xhigh, so this is not an apples-to-apples regressionsource
- FrontierCode v1.1 Extended63.6%medium reasoning; older Fable 5 is 64.9 at xhigh; the card says some out-of-scope edits are counted as failuressource
- FrontierSWE v20.57 dominance scoreProximal agent harness; max reasoning; 5 trials per tasksource
- Terminal-Bench 4.055.8%Claude Code with --bare; max thinking; 15 trials per tasksource
- Terminal-Bench-Science 0.152.6%Claude Code with --bare; max thinking; 10 trials per task, 700 trials totalsource
- CursorBench 3.2.073.4%max effort; 5 runs; Cursor independently reports the benchmark, but no steward check was foundsource
- OSWorld 2.077.9 % partial passproduction safeguards active, with Opus 4.8 fallback; default 1080p; max 500 actions; Pass@1, average of 5 runssource
- OSWorld 2.041.7 % strict passsame production-safeguard setup as the partial score; results supersede older OSWorld figures and are not comparable to earlier task releasessource
- OfficeQA Pro69%Messages API; extracted text sandbox; code execution; output limit 128K; production safeguards and fallbacksource
- Legal Agent Benchmark19.09 % all-passpublic Messages API; adaptive max; internal reimplementation; 1,235 problems, 16 defects excluded; n=5source
- Legal Agent Benchmark90.81 % criterion passsame internal benchmark and production-safeguard setup; criterion pass is not the all-pass metricsource
- GDPval-AA v21853 EloArtificial Analysis; max effort; shell and web; blind pairwise comparison over 220 tasks and 44 occupationssource
- Toolathlon Verified77.8 % Pass@1108 tasks; adaptive max; 3 trials each; production safety active with Opus 4.8 fallbacksource
- Toolathlon Verified81.5 % Pass@3same 108-task internal harness; three attemptssource
- AutomationBench31.4%max effort; private held-out leaderboard; production safeguardssource
- ARC-AGI-197.5%max effort; semi-private setsource
- ARC-AGI-290%max effort; semi-private setsource
- ARC-AGI-197.5%ARC Prize's semi-private evaluation; max effortsource
- ARC-AGI-290%ARC Prize's semi-private evaluation; max effortsource
- Maximum output128000 tokensmodel recordsource
- FrontierMath v2 — Tiers 1-30.9 scoreFable 5.1 max run; verification_code scorersource
- FrontierMath v2 — Tier 40.88 scoreFable 5.1 max run; verification_code scorersource
Claude Mythos 5.1
Anthropic2026-09closed (proprietary)
6 claimed · none independently checked
- Terminal-Bench 4.060.9%Claude Code with --bare; max thinking; 10 trials per task; Mythos safeguards configurationsource
- BioMysteryBench Human Solvable90.3%bash and file editor; restricted domains; adaptive max; Mythos access configurationsource
- BioMysteryBench Human Difficult44.1%same BioMysteryBench tool setup; difficult subsetsource
- SpatialBench Verified77.6%life-science evaluation; tool and grader details as specified in the system cardsource
- ProteinGym Hard49.3%maximum adaptive reasoning; no external tools; protein-design subsetsource
- Protocol Troubleshooting70.2%bash, file editor, and web search for protocols; maximum adaptive reasoningsource
Gemini 3.8 Flash
Google DeepMind2026-09closed (proprietary)
2 claimed · none independently checked
Gemini 3.7 Flash
Google DeepMind2026-08closed (proprietary)
9 claimed · none independently checked
- Artificial Analysis Intelligence Index56 index pointsGoogle comparison table; Aug 2026 snapshot; composition and settings not statedsource
- FrontierCode 1.1 Main43.6%Google comparison table; production code-quality setting; harness not statedsource
- DeepSWE v1.165.3%Google comparison table; long-horizon engineering; harness not statedsource
- Code Arena1588 EloGoogle comparison table; arena protocol not statedsource
- Terminal-Bench 2.185.8%Google comparison table; agentic terminal coding; harness not statedsource
- AutomationBench30.4%Google comparison table; private set; harness not statedsource
- GDPval-AA v21525 EloArtificial Analysis comparison row; task and grader details not stated in the cardsource
- CharXiv (no tools)84.5%no tools, as named in Google's tablesource
- CharXiv (with tools)88.7%tools on, as named in Google's tablesource
GPT-5.6 Cyber
OpenAI2026-08closed (proprietary)
0 claimed · none independently checked
Qwen3.8-Flash-Next
Qwen2026-08open weights (Qwen Community License 1.0)
7 claimed · none independently checked
- DeepSWE v1.158.7%highest across Claude Code and mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
- SWE-bench Pro62.5%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
- SWE-bench Multilingual81%mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
- Toolathlon Verified73.5 % Pass@1Qwen model-card comparison table; default thinking mode; exact harness not statedsource
- GPQA Diamond91.7%Qwen model-card comparison table; default thinking mode; exact shot count not statedsource
- HLE35.9%Qwen model-card comparison table; judged by GPT-4o; exact subset and harness not statedsource
- LiveCodeBench v691.9%Qwen model-card comparison table; exact harness and date window not statedsource
Every model we could source, newest first
68 models, 175 claims, every one carrying a reachable primary source. Models with no sourced number are absent by design — an announcement is not a result.All 68 models and all 175 numbers, each linked to where it was published. If a company announced a model but published no results, it is not here.
Claude Fable 5.1
Anthropic2026-09closed (proprietary)
26 claimed · 4 independently checked
- SWE-bench Pro81.2 % resolvedadaptive thinking, max effort; average of 5 trialssource
- SWE-bench Multilingual89.1 % resolved300 problems across 9 languages; adaptive thinking max; average of 5 trialssource
- SWE-bench Multimodal54.7 % resolvedadaptive thinking max; average of 5 trialssource
- DeepSWE v1.167.4%113 tasks; adaptive thinking max; average of 5 trials; hidden tests can reject valid but non-reference solutionssource
- FrontierCode v1.1 Main50.9%medium reasoning; Anthropic comparison table; older Fable 5 is 53.5 at xhigh, so this is not an apples-to-apples regressionsource
- FrontierCode v1.1 Extended63.6%medium reasoning; older Fable 5 is 64.9 at xhigh; the card says some out-of-scope edits are counted as failuressource
- FrontierSWE v20.57 dominance scoreProximal agent harness; max reasoning; 5 trials per tasksource
- Terminal-Bench 4.055.8%Claude Code with --bare; max thinking; 15 trials per tasksource
- Terminal-Bench-Science 0.152.6%Claude Code with --bare; max thinking; 10 trials per task, 700 trials totalsource
- CursorBench 3.2.073.4%max effort; 5 runs; Cursor independently reports the benchmark, but no steward check was foundsource
- OSWorld 2.077.9 % partial passproduction safeguards active, with Opus 4.8 fallback; default 1080p; max 500 actions; Pass@1, average of 5 runssource
- OSWorld 2.041.7 % strict passsame production-safeguard setup as the partial score; results supersede older OSWorld figures and are not comparable to earlier task releasessource
- OfficeQA Pro69%Messages API; extracted text sandbox; code execution; output limit 128K; production safeguards and fallbacksource
- Legal Agent Benchmark19.09 % all-passpublic Messages API; adaptive max; internal reimplementation; 1,235 problems, 16 defects excluded; n=5source
- Legal Agent Benchmark90.81 % criterion passsame internal benchmark and production-safeguard setup; criterion pass is not the all-pass metricsource
- GDPval-AA v21853 EloArtificial Analysis; max effort; shell and web; blind pairwise comparison over 220 tasks and 44 occupationssource
- Toolathlon Verified77.8 % Pass@1108 tasks; adaptive max; 3 trials each; production safety active with Opus 4.8 fallbacksource
- Toolathlon Verified81.5 % Pass@3same 108-task internal harness; three attemptssource
- AutomationBench31.4%max effort; private held-out leaderboard; production safeguardssource
- ARC-AGI-197.5%max effort; semi-private setsource
- ARC-AGI-290%max effort; semi-private setsource
- ARC-AGI-197.5%ARC Prize's semi-private evaluation; max effortsource
- ARC-AGI-290%ARC Prize's semi-private evaluation; max effortsource
- Maximum output128000 tokensmodel recordsource
- FrontierMath v2 — Tiers 1-30.9 scoreFable 5.1 max run; verification_code scorersource
- FrontierMath v2 — Tier 40.88 scoreFable 5.1 max run; verification_code scorersource
Claude Mythos 5.1
Anthropic2026-09closed (proprietary)
6 claimed · none independently checked
- Terminal-Bench 4.060.9%Claude Code with --bare; max thinking; 10 trials per task; Mythos safeguards configurationsource
- BioMysteryBench Human Solvable90.3%bash and file editor; restricted domains; adaptive max; Mythos access configurationsource
- BioMysteryBench Human Difficult44.1%same BioMysteryBench tool setup; difficult subsetsource
- SpatialBench Verified77.6%life-science evaluation; tool and grader details as specified in the system cardsource
- ProteinGym Hard49.3%maximum adaptive reasoning; no external tools; protein-design subsetsource
- Protocol Troubleshooting70.2%bash, file editor, and web search for protocols; maximum adaptive reasoningsource
Gemini 3.8 Flash
Google DeepMind2026-09closed (proprietary)
2 claimed · none independently checked
Gemini 3.7 Flash
Google DeepMind2026-08closed (proprietary)
9 claimed · none independently checked
- Artificial Analysis Intelligence Index56 index pointsGoogle comparison table; Aug 2026 snapshot; composition and settings not statedsource
- FrontierCode 1.1 Main43.6%Google comparison table; production code-quality setting; harness not statedsource
- DeepSWE v1.165.3%Google comparison table; long-horizon engineering; harness not statedsource
- Code Arena1588 EloGoogle comparison table; arena protocol not statedsource
- Terminal-Bench 2.185.8%Google comparison table; agentic terminal coding; harness not statedsource
- AutomationBench30.4%Google comparison table; private set; harness not statedsource
- GDPval-AA v21525 EloArtificial Analysis comparison row; task and grader details not stated in the cardsource
- CharXiv (no tools)84.5%no tools, as named in Google's tablesource
- CharXiv (with tools)88.7%tools on, as named in Google's tablesource
GPT-5.6 Cyber
OpenAI2026-08closed (proprietary)
0 claimed · none independently checked
Qwen3.8-Flash-Next
Qwen2026-08open weights (Qwen Community License 1.0)
7 claimed · none independently checked
- DeepSWE v1.158.7%highest across Claude Code and mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
- SWE-bench Pro62.5%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
- SWE-bench Multilingual81%mini-SWE-agent; temperature 1.0, top_p 0.95, 256K contextsource
- Toolathlon Verified73.5 % Pass@1Qwen model-card comparison table; default thinking mode; exact harness not statedsource
- GPQA Diamond91.7%Qwen model-card comparison table; default thinking mode; exact shot count not statedsource
- HLE35.9%Qwen model-card comparison table; judged by GPT-4o; exact subset and harness not statedsource
- LiveCodeBench v691.9%Qwen model-card comparison table; exact harness and date window not statedsource
Qwen3.8-27B
Qwen2026-08open weights (Apache License 2.0)
6 claimed · none independently checked
- Terminal Bench 2.173%Terminus harness; model-card comparison table; other settings not statedsource
- SWE-bench Pro61.7%Claude Code; temperature 1.0, top_p 0.95, 256K context; corrected tasks and re-evaluated baselinessource
- DeepSWE v1.142.2%Claude Code; temperature 1.0, top_p 1.0; 256K contextsource
- OSWorld-Verified84.3%OSWorld scaffold; model-card table; tool and action settings not statedsource
- HLE30.8%model-card comparison table; judged by GPT-4o; exact subset not statedsource
- LiveCodeBench v690.3%model-card comparison table; exact harness and date window not statedsource
GLM-5.3
Zhipu AI / Z.ai2026-08open weights (glm-5.3)
8 claimed · none independently checked
- Terminal Bench 2.188.2%Claude Code 2.1.207; temperature 1.0, top_p 1.0, max_new_tokens 65,536; 6-hour timeoutsource
- Terminal Bench 3.028.3%Claude Code 2.1.207; max reasoning; 400K context; 128K output; avg@3; Tool Search disabled; official task verifiersource
- DeepSWE v1.166.9%mini-SWE-agent; temperature 0.95, top_p 1.0; 6-hour timeout; 400K contextsource
- FrontierSWE78.1 dominance scoreProximal evaluation; max effort; 1M context; 128K output; score current as of 2026-08-14source
- SWE-Marathon v1.142.5%Claude Code 2.1.207; max reasoning; temperature 1.0, top_p 0.95; 1M context; 128K outputsource
- Toolathlon Verified73 % Pass@1official evaluation service; average over 3 independent runssource
- HLE with tools62.5%temperature 1.0, top_p 0.95; max generation 163,840; max context 300K with context management; GPT-5.6 Luna medium judgesource
- GDPval-AA v21769 EloArtificial Analysis evaluation; exact task/run details not stated in cardsource
Grok 4.5
xAI2026-08closed
1 claimed · 1 independently checked
- ARC-AGI-252.6%ARC Prize independent evaluationsource
Grok 4.6
xAI2026-08closed
1 claimed · 1 independently checked
- ARC-AGI-267.1%ARC Prize independent evaluationsource
Claude Opus 5
Anthropic2026-07closed (proprietary)
3 claimed · 3 independently checked
Gemini 3.6 Flash
Google DeepMind2026-07closed (proprietary)
6 claimed · none independently checked
- SWE-Bench Pro Public58.7%Google comparison table; public split; harness not statedsource
- DeepSWE v1.149%Google comparison table; harness not statedsource
- Terminal-Bench 2.178%Terminus-2 harnesssource
- OSWorld-Verified83%Google comparison table; computer-use setting not otherwise statedsource
- CharXiv (no tools)85.2%no tools, as named in Google's tablesource
- CharXiv (with tools)89.4%tools on, as named in Google's tablesource
GPT-5.6 Sol
OpenAI2026-07closed (proprietary)
3 claimed · 3 independently checked
GPT-5.6 Terra
OpenAI2026-07closed (proprietary)
3 claimed · 3 independently checked
GPT-5.6 Luna
OpenAI2026-07closed (proprietary)
3 claimed · 3 independently checked
Kimi K3
Moonshot AI2026-07open weights (Kimi K3 License)
9 claimed · none independently checked
- GPQA Diamond93.5%max reasoning; temperature 1.0; top_p 0.95 for single-step taskssource
- HLE-Full43.5%max reasoning; without tools; temperature 1.0; top_p 0.95 for single-step taskssource
- HLE-Full56%max reasoning; with tools; general tools; temperature 1.0; top_p 0.95 single-step and 1.0 agenticsource
- DeepSWE v1.167.5%Kimi Code harness; the card says the 67.3 leaderboard result is mini-SWE-agent, but reports 67.5 for K3's runsource
- Terminal-Bench 2.188.3%Kimi Code harness; max reasoning; agentic top_p 1.0source
- FrontierSWE81.2 dominance scoreKimi Code harness; dominance recomputed with official script; current as of 2026-07-16source
- SWE-Marathon42%Claude Code harness; H20-calibrated branch of tasks as of 2026-07-09, before final v1.1source
- GDPval-AA v21686 EloArtificial Analysis scores cited by Kimi; board observed as of 2026-07-23source
- Toolathlon-Verified76.5 % Pass@1Kimi comparison table; grader and trial count not stated in the rowsource
Claude Sonnet 5
Anthropic2026-06closed (proprietary)
0 claimed · none independently checked
MiniMax-M3
MiniMax2026-06open weights (MiniMax Community License)
0 claimed · none independently checked
Claude Opus 4.8
Anthropic2026-05closed (proprietary)
0 claimed · none independently checked
Gemini 3.5 Flash
Google DeepMind2026-05closed (proprietary)
5 claimed · none independently checked
- Terminal-Bench 2.176.2%Terminus-2 harnesssource
- SWE-Bench Pro Public55.1%single attempt, public splitsource
- MCP Atlas83.6%Google comparison table; subset/harness not statedsource
- Toolathlon56.5%Google comparison table; attempt count not statedsource
- OSWorld-Verified78.4%Google comparison table; computer-use setting not statedsource
Gemma 4 31B
Google DeepMind2026-04open weights (Apache License 2.0)
9 claimed · none independently checked
- MMLU Pro85.2%instruction-tuned model card result; other harness details not statedsource
- AIME 202689.2%no toolssource
- LiveCodeBench v680%instruction-tuned model card result; harness not statedsource
- Codeforces2150 Elomodel card result; contest protocol not statedsource
- GPQA Diamond84.3%instruction-tuned model card result; tools not statedsource
- τ276.9%average over 3source
- HLE19.5%no toolssource
- HLE26.5%with searchsource
- MRCR v2, 8-needle, 128K66.4 % average128K context; 8 needles; average metricsource
GPT-5.5
OpenAI2026-04closed (proprietary)
0 claimed · none independently checked
GPT-5.5 Pro
OpenAI2026-04closed (proprietary)
0 claimed · none independently checked
DeepSeek-V4-Pro
DeepSeek2026-04open weights (MIT License)
8 claimed · none independently checked
- MMLU90.1 % EMbase model; 5-shotsource
- MMLU-Pro73.5 % EMbase model; 5-shotsource
- LongBench-V251.5 % EMbase model; 1-shotsource
- HLE37.7 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, no tools rowsource
- HLE48.2 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; with toolssource
- Terminal Bench 2.067.9 % accuracyDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
- SWE Verified80.6 % resolvedDeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
- Toolathlon51.8 % Pass@1DeepSeek-V4-Pro-Max; max reasoning; comparison table, harness not statedsource
DeepSeek-V4-Flash
DeepSeek2026-04open weights (MIT License)
0 claimed · none independently checked
Kimi K2.6
Moonshot AI2026-04open weights (Modified MIT License)
6 claimed · none independently checked
- HLE-Full36.4%max generation 98,304; without toolssource
- HLE-Full55.5%max generation 98,304; with tools uses search, code interpreter, and web browsing plus context managementsource
- Terminal-Bench 2.066.7%Terminus-2; preserved thinking; default agent frameworksource
- SWE-Bench Pro58.6%in-house SWE-agent-derived harness with bash/createfile/insert/view/strreplace/submit toolssource
- SWE-Bench Verified80.2%same in-house minimal-tool harness; default temperature 1.0, top_p 1.0, 262,144 contextsource
- LiveCodeBench v689.6%model-card comparison table; exact sampling not stated in rowsource
MiniMax-M2.7
MiniMax2026-04research only (NON-COMMERCIAL LICENSE)
10 claimed · none independently checked
- MLE Bench Lite66.6 % medal rate22 ML competitions; card describes internal model self-evolution comparisonsource
- SWE-Pro56.22%MiniMax card; scaffold and run count not stated in the paragraphsource
- SWE-bench Multilingual76.5%MiniMax card; scaffold and run count not stated in the paragraphsource
- Multi SWE Bench52.7%MiniMax card; scaffold and run count not stated in the paragraphsource
- VIBE-Pro55.6%MiniMax card; benchmark harness not stated in the paragraphsource
- Terminal Bench 257%MiniMax card; benchmark harness not stated in the paragraphsource
- NL2Repo39.8%MiniMax card; benchmark harness not stated in the paragraphsource
- GDPval-AA1495 EloMiniMax card; claimed highest among open-weight models; board date not statedsource
- Toolathlon46.3 % accuracyMiniMax card; benchmark harness and attempt count not statedsource
- MM Claw62.7%MiniMax card; end-to-end benchmark; harness and attempt count not statedsource
Claude Mythos Preview (early)
Anthropic2026-04
1 claimed · 1 independently checked
- METR time horizon (TH1.1)1044.78 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5.4
OpenAI2026-03closed (proprietary)
1 claimed · 1 independently checked
- METR time horizon (TH1.1)341.74 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5.4 Pro
OpenAI2026-03closed (proprietary)
0 claimed · none independently checked
GPT-5.4 mini
OpenAI2026-03closed (proprietary)
0 claimed · none independently checked
Gemini 3.1 Pro
Google DeepMind2026-02closed (proprietary)
6 claimed · 1 independently checked
- HLE (no tools)44.4%no toolssource
- HLE (Search + Code)51.4%Search blocklist + Codesource
- GPQA Diamond (no tools)94.3%no toolssource
- Terminal-Bench 2.068.5%Terminus-2 harnesssource
- SWE-Bench Verified80.6%single attemptsource
- METR time horizon (TH1.1)384.15 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5.3-Codex
OpenAI2026-02closed (proprietary)
1 claimed · 1 independently checked
- METR time horizon (TH1.1)349.53 minutesp50 horizon length; METR-Horizon-v1.1source
Claude Opus 4.6
Anthropic2026-02
1 claimed · 1 independently checked
- METR time horizon (TH1.1)718.81 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5.2-Codex
OpenAI2026-01closed (proprietary)
0 claimed · none independently checked
GPT-5.1-Codex-Max
OpenAI2025-12closed (proprietary)
1 claimed · 1 independently checked
- METR time horizon (TH1.1)223.71 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5.2
OpenAI2025-12
1 claimed · 1 independently checked
- METR time horizon (TH1.1)352.25 minutesp50 horizon length; METR-Horizon-v1.1source
Devstral 2 123B
Mistral2025-12
2 claimed · 1 independently checked
Devstral Small 2 24B
Mistral2025-12open weights (Apache 2.0)
2 claimed · 1 independently checked
GPT-5.1-Codex
OpenAI2025-11closed (proprietary)
0 claimed · none independently checked
GPT-5.1-Codex Mini
OpenAI2025-11closed (proprietary)
0 claimed · none independently checked
Gemini 3 Pro
Google DeepMind2025-11
1 claimed · 1 independently checked
- METR time horizon (TH1.1)224.33 minutesp50 horizon length; METR-Horizon-v1.1source
Claude Opus 4.5
Anthropic2025-11
1 claimed · 1 independently checked
- METR time horizon (TH1.1)292.99 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5-Codex
OpenAI2025-09closed (proprietary)
0 claimed · none independently checked
Magistral Small 1.2
Mistral2025-09open weights (Apache 2.0)
2 claimed · 1 independently checked
Claude Opus 4.1
Anthropic2025-08
1 claimed · 1 independently checked
- METR time horizon (TH1.1)100.47 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-5
OpenAI2025-08
1 claimed · 1 independently checked
- METR time horizon (TH1.1)203.01 minutesp50 horizon length; METR-Horizon-v1.1source
Devstral Small 1.1
Mistral2025-07open weights (Apache 2.0)
2 claimed · 2 independently checked
Claude Opus 4
Anthropic2025-05
1 claimed · 1 independently checked
- METR time horizon (TH1.1)100.37 minutesp50 horizon length; METR-Horizon-v1.1source
o3
OpenAI2025-04closed (proprietary)
1 claimed · 1 independently checked
- METR time horizon (TH1.1)119.73 minutesp50 horizon length; METR-Horizon-v1.1source
o4-mini
OpenAI2025-04closed (proprietary)
0 claimed · none independently checked
Llama 4 Maverick
Meta2025-04
1 claimed · 1 independently checked
- SWE-bench Verified21.04%checked: truesource
Llama 4 Scout
Meta2025-04
1 claimed · 1 independently checked
- SWE-bench Verified9.06%checked: truesource
Mistral Small 3.1
Mistral2025-03open weights (Apache 2.0)
2 claimed · 1 independently checked
Claude 3.7 Sonnet
Anthropic2025-02
1 claimed · 1 independently checked
- METR time horizon (TH1.1)60.39 minutesp50 horizon length; METR-Horizon-v1.1source
Grok 3
xAI2025-02closed
1 claimed · 1 independently checked
- ARC-AGI-20%ARC Prize independent evaluationsource
o1
OpenAI2024-12
1 claimed · 1 independently checked
- METR time horizon (TH1.1)38.83 minutesp50 horizon length; METR-Horizon-v1.1source
Claude 3.5 Sonnet (Oct 2024)
Anthropic2024-10
1 claimed · 1 independently checked
- METR time horizon (TH1.1)20.52 minutesp50 horizon length; METR-Horizon-v1.1source
o1-preview
OpenAI2024-09
1 claimed · 1 independently checked
- METR time horizon (TH1.1)20.33 minutesp50 horizon length; METR-Horizon-v1.1source
Claude 3.5 Sonnet (Jun 2024)
Anthropic2024-06
1 claimed · 1 independently checked
- METR time horizon (TH1.1)11.4 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-4o
OpenAI2024-05
1 claimed · 1 independently checked
- METR time horizon (TH1.1)6.99 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-4 Turbo
OpenAI2024-04
1 claimed · 1 independently checked
- METR time horizon (TH1.1)3.73 minutesp50 horizon length; METR-Horizon-v1.1source
Claude 3 Opus
Anthropic2024-03
1 claimed · 1 independently checked
- METR time horizon (TH1.1)3.95 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-4 (1106)
OpenAI2023-11
1 claimed · 1 independently checked
- METR time horizon (TH1.1)4.04 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-4
OpenAI2023-03
1 claimed · 1 independently checked
- METR time horizon (TH1.1)3.99 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-3.5 Turbo Instruct
OpenAI2022-03
1 claimed · 1 independently checked
- METR time horizon (TH1.1)0.6 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-3 (davinci-002)
OpenAI2020-05
1 claimed · 1 independently checked
- METR time horizon (TH1.1)0.14 minutesp50 horizon length; METR-Horizon-v1.1source
GPT-2
OpenAI2019-02
1 claimed · 1 independently checked
- METR time horizon (TH1.1)0.05 minutesp50 horizon length; METR-Horizon-v1.1source
What counts as checked, and what we left out
Every number was read from a primary source — a lab's own model card, system card or pricing page, or a steward's published results. Nothing came from an aggregator or a news write-up. A result counts as checkedonly when a benchmark's own steward published or re-ran it, and that steward is always named. Numbers a lab published only as an image or a chart are recorded as unreadable rather than guessed at, and models with no sourced figure are omitted entirely. Observed 2026-09-02; the full working, including everything we could not verify, is in the repository.