LLM benchmark leaderboards
EvalsAs of August 20, 2026, this index ranks 696 public LLM evals next to live API token prices. Every score on a model or comparison page links here. Official boards own the methodology; these pages add $/M.
| Benchmark | US searches / mo | Models scored | What it measures |
|---|---|---|---|
| HLE | 6,600 | 104 | Humanity's Last Exam: 2,500 expert questions designed so search engines fail. |
| SWE Bench | 6,600 | 8 | the SWE Bench score |
| Humanity's Last Exam (No Tools) | 6,600 | 5 | the Humanity's Last Exam (No Tools) score |
| Humanity's Last Exam (With Tools) | 6,600 | 5 | the Humanity's Last Exam (With Tools) score |
| Terminal-Bench 2.0 | 5,400 | 68 | the Terminal-Bench 2.0 score |
| Terminal-Bench | 5,400 | 67 | Terminal-Bench: can the model finish jobs in a real shell. |
| Terminal-Bench 2.1 | 5,400 | 52 | Terminal-Bench: can the model finish jobs in a real shell. |
| Terminal-Bench 2 | 5,400 | 45 | the Terminal-Bench 2 score |
| Terminal-Bench 3.0 | 5,400 | 3 | the Terminal-Bench 3.0 score |
| ARC-AGI-1 Verified | 3,600 | 72 | the ARC-AGI-1 Verified score |
| Chatbot Arena | 3,600 | 16 | Chatbot Arena human votes: people pick the better reply in a blind test. That is a writing and taste rank, not a trivia quiz. |
| ARC-AGI-3 | 3,600 | 10 | the ARC-AGI-3 score |
| Arc Agi | 3,600 | 8 | the Arc Agi score |
| DeepSWE | 2,900 | 34 | DeepSWE: repository-level coding agents on real GitHub work. |
| SWE-bench Verified | 2,400 | 137 | the SWE-bench Verified score |
| SWE-bench Pro | 2,400 | 52 | the SWE-bench Pro score |
| SWE-bench Pro (Public) | 2,400 | 10 | the SWE-bench Pro (Public) score |
| MMLU | 1,900 | 87 | MMLU, a multiple-choice knowledge exam. It does not measure prose. |
| ARC-AGI 2 | 1,900 | 69 | ARC-AGI: novel visual puzzles, not memorized exams. |
| LiveBench | 1,900 | 55 | the LiveBench score |
| OSWorld | 1,900 | 38 | OSWorld: can the model drive a real desktop (click, type, finish the task). |
| OSWorld Verified | 1,900 | 36 | OSWorld-Verified: cleaned desktop computer-use tasks with an execution check. |
| GDPval | 1,900 | 16 | GDPval: professional knowledge-work, closer to office output than trivia. |
| OSWorld 2.0 | 1,900 | 12 | the OSWorld 2.0 score |
| Screenspot | 1,900 | 11 | the Screenspot score |
| ARC-AGI-2 Verified | 1,900 | 2 | the ARC-AGI-2 Verified score |
| ARC-AGI 2 | 1,900 | 1 | the ARC-AGI 2 score |
| SimpleBench | 1,600 | 59 | the SimpleBench score |
| GSM8K | 1,600 | 36 | the GSM8K score |
| Treebench | 1,600 | 1 | the Treebench score |
| LiveCodeBench | 1,300 | 128 | LiveCodeBench: contest programming on problems released after training. |
| tau-bench Retail | 1,300 | 44 | tau-bench: multi-turn tool use on customer-support workflows. |
| LiveCodeBench v6 | 1,300 | 43 | the LiveCodeBench v6 score |
| tau2-bench (Telecom) | 1,300 | 31 | the tau2-bench (Telecom) score |
| tau2-bench Telecom | 1,300 | 30 | the tau2-bench Telecom score |
| Tau Bench | 1,300 | 6 | the Tau Bench score |
| GPQA | 1,000 | 189 | GPQA: graduate-level science questions that resist simple search. |
| BrowseComp | 1,000 | 59 | the BrowseComp score |
| Tvbench | 1,000 | 2 | the Tvbench score |
| Deck Bench | 1,000 | 1 | the Deck Bench score |
| AIME 2025 | 880 | 170 | AIME 2025: American invitational math contest problems. |
| GPQA Diamond | 880 | 126 | GPQA Diamond: graduate-level science questions written so Google cannot solve them. |
| Aime 2026 | 880 | 14 | the Aime 2026 score |
| HealthBench | 880 | 9 | HealthBench: medical conversation quality, not multiple-choice science. |
| MMLU-Pro | 720 | 165 | MMLU-Pro, a hard multiple-choice knowledge quiz. It is not a writing test. |
| HumanEval | 720 | 51 | the HumanEval score |
| MMMU | 590 | 98 | MMMU: college-level questions on diagrams, charts, and images. |
| FrontierMath | 590 | 53 | FrontierMath: research-level math that stays hard after models ace AIME. |
| CursorBench | 590 | 14 | the CursorBench score |
| Big Bench | 590 | 2 | the Big Bench score |
| Mle Bench | 590 | 2 | the Mle Bench score |
| SimpleQA | 480 | 86 | SimpleQA, which checks whether the model makes up facts. |
| IFEval | 480 | 49 | IFEval, which checks whether the model followed the prompt (format, length, constraints). That is instruction-following, not 'sounds good.' |
| MATH 500 | 480 | 42 | MATH-500: textbook and contest math problems. |
| SkillsBench | 480 | 30 | SkillsBench: can the agent actually use bundled skills and tools on multi-step tasks, not just chat about them. |
| HellaSwag | 480 | 23 | HellaSwag, a sentence-completion quiz, not a writing sample. |
| Mt Bench | 480 | 8 | the Mt Bench score |
| Omnidocbench | 480 | 2 | the Omnidocbench score |
| AIME 2024 | 390 | 41 | the AIME 2024 score |
| Vending Bench 2 | 390 | 4 | the Vending Bench 2 score |
| Pinchbench | 390 | 3 | the Pinchbench score |
| Eq Bench | 390 | 2 | the Eq Bench score |
| MMMU-Pro | 320 | 54 | MMMU-Pro: a harder multimodal exam than MMMU. |
| IFBench | 320 | 25 | IFBench, a harder instruction-following set than IFEval. |
| ScreenSpot Pro | 320 | 21 | the ScreenSpot Pro score |
| Mrcr | 320 | 7 | the Mrcr score |
| Paperbench | 320 | 3 | the Paperbench score |
| Cybench | 320 | 2 | the Cybench score |
| DeepResearch Bench | 260 | 20 | the DeepResearch Bench score |
| Triviaqa | 260 | 7 | the Triviaqa score |
| Gdpval Aa | 260 | 6 | the Gdpval Aa score |
| Mmbench | 260 | 2 | the Mmbench score |
| TriviaQA | 260 | 2 | the TriviaQA score |
| SciCode | 210 | 90 | the SciCode score |
| ToolAthlon | 210 | 37 | the ToolAthlon score |
| MathVista | 210 | 30 | the MathVista score |
| Ocrbench | 210 | 16 | the Ocrbench score |
| Longbench V2 | 210 | 12 | the Longbench V2 score |
| Program Bench | 210 | 5 | the Program Bench score |
| Livecodebench Pro | 210 | 3 | the Livecodebench Pro score |
| Cl Bench | 210 | 2 | the Cl Bench score |
| Bixbench | 210 | 1 | the Bixbench score |
| MMMLU | 170 | 38 | the MMMLU score |
| SuperGPQA | 170 | 28 | the SuperGPQA score |
| Mvbench | 170 | 13 | the Mvbench score |
| Claw Eval | 170 | 11 | the Claw Eval score |
| Exploitbench | 170 | 5 | the Exploitbench score |
| Posttrainbench | 170 | 5 | the Posttrainbench score |
| Swe Marathon | 170 | 4 | the Swe Marathon score |
| SWE-bench Multilingual | 140 | 38 | the SWE-bench Multilingual score |
| MRCR v2 | 140 | 22 | MRCR: can the model still find planted facts inside a long prompt. |
| LVBench | 140 | 21 | the LVBench score |
| WinoGrande | 140 | 16 | the WinoGrande score |
| Longvideobench | 140 | 2 | the Longvideobench score |
| Evalplus | 140 | 1 | the Evalplus score |
| FrontierMath Tier 4 | 110 | 40 | the FrontierMath Tier 4 score |
| Fiction.liveBench | 110 | 35 | the Fiction.liveBench score |
| Big Bench Hard | 110 | 13 | the Big Bench Hard score |
| BBH | 110 | 12 | the BBH score |
| HealthBench Professional | 110 | 9 | the HealthBench Professional score |
| Finance Agent | 110 | 8 | the Finance Agent score |
| Zerobench | 110 | 7 | the Zerobench score |
| Multi Swe Bench | 110 | 6 | the Multi Swe Bench score |
| Biomysterybench | 110 | 3 | the Biomysterybench score |
| Ocrbench V2 | 110 | 3 | the Ocrbench V2 score |
| Swe Lancer | 110 | 3 | the Swe Lancer score |
| Deep Swe | 110 | 1 | the Deep Swe score |
| Simpleqa Verified | 110 | 1 | the Simpleqa Verified score |
| Swe Atlas | 110 | 1 | the Swe Atlas score |
| MMLU-Redux | 90 | 37 | the MMLU-Redux score |
| Legal Agent Benchmark | 90 | 14 | the Legal Agent Benchmark score |
| Writingbench | 90 | 14 | the Writingbench score |
| Mmlongbench Doc | 90 | 5 | the Mmlongbench Doc score |
| Tau3 Bench | 90 | 5 | the Tau3 Bench score |
| Swe Bench Multimodal | 90 | 4 | the Swe Bench Multimodal score |
| Wildclawbench | 90 | 4 | the Wildclawbench score |
| Acebench | 90 | 2 | the Acebench score |
| Terminal Bench Hard | 90 | 2 | the Terminal Bench Hard score |
| Web Bench | 90 | 1 | the Web Bench score |
| ProofBench | 70 | 52 | the ProofBench score |
| CyberBench | 70 | 35 | the CyberBench score |
| VideoMMMU | 70 | 25 | the VideoMMMU score |
| T2 Bench | 70 | 15 | the T2 Bench score |
| AutomationBench | 70 | 13 | the AutomationBench score |
| C Eval | 70 | 11 | the C Eval score |
| Lifescibench | 70 | 3 | the Lifescibench score |
| Cfeval | 70 | 2 | the Cfeval score |
| Genebench | 70 | 2 | the Genebench score |
| Graphwalks | 70 | 2 | the Graphwalks score |
| Automation Bench | 70 | 1 | the Automation Bench score |
| Big Bench Audio | 70 | 1 | the Big Bench Audio score |
| Swt Bench | 70 | 1 | the Swt Bench score |
| Webdev Arena | 70 | 1 | the Webdev Arena score |
| Worldbench | 70 | 1 | the Worldbench score |
| APEX-Agents | 50 | 50 | the APEX-Agents score |
| Browsecomp Zh | 50 | 12 | the Browsecomp Zh score |
| HealthBench Hard | 50 | 8 | the HealthBench Hard score |
| Agieval | 50 | 4 | the Agieval score |
| Job Bench | 50 | 4 | the Job Bench score |
| Blueprint Bench 2 | 50 | 3 | the Blueprint Bench 2 score |
| Genebench Pro | 50 | 3 | the Genebench Pro score |
| Global Mmlu | 50 | 2 | the Global Mmlu score |
| Motionbench | 50 | 2 | the Motionbench score |
| Labbench2 | 50 | 1 | the Labbench2 score |
| Osworld G | 50 | 1 | the Osworld G score |
| Ovobench | 50 | 1 | the Ovobench score |
| Phybench | 50 | 1 | the Phybench score |
| Yc Bench | 50 | 1 | the Yc Bench score |
| Mmlu Prox | 40 | 25 | the Mmlu Prox score |
| Muirbench | 40 | 9 | the Muirbench score |
| Ojbench | 40 | 9 | the Ojbench score |
| Complexfuncbench | 40 | 6 | the Complexfuncbench score |
| Countbench | 40 | 6 | the Countbench score |
| Big Bench Extra Hard | 40 | 5 | the Big Bench Extra Hard score |
| Benchcad | 40 | 4 | the Benchcad score |
| Cmmlu | 40 | 4 | the Cmmlu score |
| Zclawbench | 40 | 4 | the Zclawbench score |
| Gdpval Aa V2 | 40 | 2 | the Gdpval Aa V2 score |
| Bankertoolbench | 40 | 1 | the Bankertoolbench score |
| Posttrain Bench | 40 | 1 | the Posttrain Bench score |
| Vibe Code Bench | 30 | 59 | the Vibe Code Bench score |
| MCP Atlas | 30 | 37 | the MCP Atlas score |
| Imo Answerbench | 30 | 20 | the Imo Answerbench score |
| AA-LCR | 30 | 15 | long-context retrieval: can the model use a huge window, not just advertise one. |
| Vita Bench | 30 | 9 | the Vita Bench score |
| Tir Bench | 30 | 4 | the Tir Bench score |
| Sec Bench Pro | 30 | 3 | the Sec Bench Pro score |
| Workspace Bench | 30 | 2 | the Workspace Bench score |
| Cmt Benchmark | 30 | 1 | the Cmt Benchmark score |
| Frontier-Bench v0.1 | 30 | 1 | the Frontier-Bench v0.1 score |
| Frontier Swe | 30 | 1 | the Frontier Swe score |
| Gdpval Aa Elo | 30 | 1 | the Gdpval Aa Elo score |
| Harvey Lab | 30 | 1 | the Harvey Lab score |
| Livesqlbench | 30 | 1 | the Livesqlbench score |
| Longcodebench | 30 | 1 | the Longcodebench score |
| Mle Bench Lite | 30 | 1 | the Mle Bench Lite score |
| Profbench | 30 | 1 | the Profbench score |
| Qwenclawbench | 30 | 1 | the Qwenclawbench score |
| Swe Perf | 30 | 1 | the Swe Perf score |
| DeepSWE 1.1 | 20 | 27 | DeepSWE 1.1: repository-level coding agents on real GitHub work. |
| GDP-PDF | 20 | 22 | the GDP-PDF score |
| Arc C | 20 | 21 | the Arc C score |
| Tau Bench Airline | 20 | 21 | the Tau Bench Airline score |
| Arena-Hard | 20 | 20 | the Arena-Hard score |
| Bfcl V3 | 20 | 17 | the Bfcl V3 score |
| Video Mme | 20 | 15 | the Video Mme score |
| Hallusion Bench | 20 | 14 | the Hallusion Bench score |
| Global Mmlu Lite | 20 | 10 | the Global Mmlu Lite score |
| Vibe Eval | 20 | 8 | the Vibe Eval score |
| Video-MME | 20 | 7 | the Video-MME score |
| Androidworld | 20 | 5 | the Androidworld score |
| Graphwalks BFS (256K-1M) | 20 | 5 | the Graphwalks BFS (256K-1M) score |
| Livecodebench V5 | 20 | 5 | the Livecodebench V5 score |
| FrontierMath (Tiers 1-3) | 20 | 4 | the FrontierMath (Tiers 1-3) score |
| Alignbench | 20 | 3 | the Alignbench score |
| Big Finance Bench | 20 | 3 | the Big Finance Bench score |
| Lmarena Text | 20 | 2 | the Lmarena Text score |
| Multilingual Mmlu | 20 | 2 | the Multilingual Mmlu score |
| Visualwebbench | 20 | 2 | the Visualwebbench score |
| Browsecomp Vl | 20 | 1 | the Browsecomp Vl score |
| Deepswe V1 1 | 20 | 1 | the Deepswe V1 1 score |
| Humaneval Plus | 20 | 1 | the Humaneval Plus score |
| Kernelbench Hard | 20 | 1 | the Kernelbench Hard score |
| Mmbench Video | 20 | 1 | the Mmbench Video score |
| Octocodingbench | 20 | 1 | the Octocodingbench score |
| Presentbench | 20 | 1 | the Presentbench score |
| Researchclawbench | 20 | 1 | the Researchclawbench score |
| Svg Bench | 20 | 1 | the Svg Bench score |
| Swe Fficiency | 20 | 1 | the Swe Fficiency score |
| HMMT 2025 | 10 | 28 | the HMMT 2025 score |
| tau2-bench Retail | 10 | 24 | the tau2-bench Retail score |
| tau2-bench Airline | 10 | 21 | the tau2-bench Airline score |
| Mathvista Mini | 10 | 19 | the Mathvista Mini score |
| Realworldqa | 10 | 19 | the Realworldqa score |
| Multi If | 10 | 18 | the Multi If score |
| Cc Ocr | 10 | 16 | the Cc Ocr score |
| Mm Mt Bench | 10 | 15 | the Mm Mt Bench score |
| Graphwalks BFS (0K-128K) | 10 | 14 | the Graphwalks BFS (0K-128K) score |
| Arena Hard V2 | 10 | 13 | the Arena Hard V2 score |
| Livebench 20241125 | 10 | 13 | the Livebench 20241125 score |
| Mmbench V1 1 | 10 | 13 | the Mmbench V1 1 score |
| Creative Writing V3 | 10 | 12 | the Creative Writing V3 score |
| Facts Grounding | 10 | 12 | the Facts Grounding score |
| Mmmu Val | 10 | 10 | the Mmmu Val score |
| Agents' Last Exam | 10 | 9 | the Agents' Last Exam score |
| Bfcl V4 | 10 | 8 | the Bfcl V4 score |
| Deep Planning | 10 | 8 | the Deep Planning score |
| Multipl E | 10 | 8 | the Multipl E score |
| Nova 63 | 10 | 8 | the Nova 63 score |
| The Agent Company | 10 | 8 | the The Agent Company score |
| Wild Bench | 10 | 8 | the Wild Bench score |
| Embspatialbench | 10 | 7 | the Embspatialbench score |
| Matharena Apex | 10 | 7 | the Matharena Apex score |
| OfficeQA Pro | 10 | 7 | the OfficeQA Pro score |
| Refspatialbench | 10 | 6 | the Refspatialbench score |
| Seal 0 | 10 | 6 | the Seal 0 score |
| Coworkbench | 10 | 4 | the Coworkbench score |
| HealthBench Consensus | 10 | 4 | the HealthBench Consensus score |
| Mrcr 1m | 10 | 4 | the Mrcr 1m score |
| Aa Briefcase | 10 | 3 | the Aa Briefcase score |
| Api Bank | 10 | 3 | the Api Bank score |
| Arc E | 10 | 3 | the Arc E score |
| Cnmo 2024 | 10 | 3 | the Cnmo 2024 score |
| Frontiermath Tier 4 V2 | 10 | 3 | the Frontiermath Tier 4 V2 score |
| Gdpval Mm | 10 | 3 | the Gdpval Mm score |
| Gorilla Benchmark Api Bench | 10 | 3 | the Gorilla Benchmark Api Bench score |
| Medchembench | 10 | 3 | the Medchembench score |
| Mls Bench Lite | 10 | 3 | the Mls Bench Lite score |
| Mmlu Cot | 10 | 3 | the Mmlu Cot score |
| Mmmuval | 10 | 3 | the Mmmuval score |
| Posttrainbench Lite | 10 | 3 | the Posttrainbench Lite score |
| Spreadsheetbench V1 | 10 | 3 | the Spreadsheetbench V1 score |
| Frontier Science Research | 10 | 2 | the Frontier Science Research score |
| Humaneval Mul | 10 | 2 | the Humaneval Mul score |
| Kimi Code Bench V2 | 10 | 2 | the Kimi Code Bench V2 score |
| Mbpp Evalplus | 10 | 2 | the Mbpp Evalplus score |
| Mimo Coding Bench | 10 | 2 | the Mimo Coding Bench score |
| Natural Questions | 10 | 2 | the Natural Questions score |
| Onemillion Bench | 10 | 2 | the Onemillion Bench score |
| Qwen Swe Bench | 10 | 2 | the Qwen Swe Bench score |
| Qwenwebbench | 10 | 2 | the Qwenwebbench score |
| Qwenworldbench | 10 | 2 | the Qwenworldbench score |
| Swe Atlas Codebase Qna | 10 | 2 | the Swe Atlas Codebase Qna score |
| Swe Bench Verified Agentic Coding | 10 | 2 | the Swe Bench Verified Agentic Coding score |
| Vibe Pro | 10 | 2 | the Vibe Pro score |
| Agent Startup Bench | 10 | 1 | the Agent Startup Bench score |
| Androidbench | 10 | 1 | the Androidbench score |
| Apex Swe | 10 | 1 | the Apex Swe score |
| Artifacts Bench | 10 | 1 | the Artifacts Bench score |
| Automationbench Aa | 10 | 1 | the Automationbench Aa score |
| Bigcodebench Hard | 10 | 1 | the Bigcodebench Hard score |
| Biolp Bench | 10 | 1 | the Biolp Bench score |
| Charxiv Reasoning | 10 | 1 | the Charxiv Reasoning score |
| Cl Bench Life | 10 | 1 | the Cl Bench Life score |
| Finance Agent V1 1 | 10 | 1 | the Finance Agent V1 1 score |
| Frontiercode Diamond | 10 | 1 | the Frontiercode Diamond score |
| Gdpval Aa V2 Elo | 10 | 1 | the Gdpval Aa V2 Elo score |
| Gdpval Rubrics | 10 | 1 | the Gdpval Rubrics score |
| Hle Full | 10 | 1 | the Hle Full score |
| Hr Bench 4k | 10 | 1 | the Hr Bench 4k score |
| Imo 2025 | 10 | 1 | the Imo 2025 score |
| Infinitebench En Mc | 10 | 1 | the Infinitebench En Mc score |
| Kernel Bench L3 | 10 | 1 | the Kernel Bench L3 score |
| Kimi Claw 24 7 Bench | 10 | 1 | the Kimi Claw 24 7 Bench score |
| Livemathematicianbench | 10 | 1 | the Livemathematicianbench score |
| Mcp Universe | 10 | 1 | the Mcp Universe score |
| Measurebench | 10 | 1 | the Measurebench score |
| Miabench | 10 | 1 | the Miabench score |
| Mm Clawbench | 10 | 1 | the Mm Clawbench score |
| Mmsibench | 10 | 1 | the Mmsibench score |
| Next.js Evals | 10 | 1 | the Next.js Evals score |
| Nih Multi Needle | 10 | 1 | the Nih Multi Needle score |
| Openai Mmlu | 10 | 1 | the Openai Mmlu score |
| Ovbench | 10 | 1 | the Ovbench score |
| Plawbench | 10 | 1 | the Plawbench score |
| Prbench Finance | 10 | 1 | the Prbench Finance score |
| Seccodebench | 10 | 1 | the Seccodebench score |
| Spreadsheetbench 2 | 10 | 1 | the Spreadsheetbench 2 score |
| Swe Atlas Test Writing | 10 | 1 | the Swe Atlas Test Writing score |
| Swe Mm | 10 | 1 | the Swe Mm score |
| Swe Review | 10 | 1 | the Swe Review score |
| Toolathlon Verified | 10 | 1 | the Toolathlon Verified score |
| Vibe V2 | 10 | 1 | the Vibe V2 score |
| Vladbench | 10 | 1 | the Vladbench score |
| Webarena Verified | 10 | 1 | the Webarena Verified score |
| Xdailybench | 10 | 1 | the Xdailybench score |
| MATH | — | 86 | the MATH score |
| Arena Code Elo | — | 84 | the Arena Code Elo score |
| CritPt | — | 84 | the CritPt score |
| Tax Eval v2 | — | 84 | the Tax Eval v2 score |
| CorpFin v2 | — | 80 | the CorpFin v2 score |
| MGSM | — | 61 | the MGSM score |
| MedCode | — | 60 | the MedCode score |
| MedScribe | — | 60 | the MedScribe score |
| SAGE | — | 55 | the SAGE score |
| CharXiv-R | — | 47 | CharXiv: can the model read scientific charts. |
| MedQA | — | 47 | the MedQA score |
| IOI | — | 45 | the IOI score |
| Alder Polyglot | — | 43 | the Alder Polyglot score |
| Code Migration | — | 43 | the Code Migration score |
| BFCL | — | 42 | the BFCL score |
| FinanceAgent v2 | — | 42 | FinanceAgent: financial analysis tasks, not a generic chat vibe. |
| EMB | — | 40 | the EMB score |
| FinanceAgent v1.1 | — | 39 | FinanceAgent: financial analysis tasks, not a generic chat vibe. |
| Arena Vision Elo | — | 37 | the Arena Vision Elo score |
| Lech Mazur Writing | — | 27 | the Lech Mazur writing eval, which scores prose quality instead of multiple-choice knowledge. |
| Ai2d | — | 25 | the Ai2d score |
| DROP | — | 24 | the DROP score |
| INCLUDE | — | 24 | the INCLUDE score |
| Arena Search Elo | — | 23 | the Arena Search Elo score |
| MathVision | — | 22 | the MathVision score |
| ERQA | — | 21 | the ERQA score |
| Hmmt25 | — | 21 | the Hmmt25 score |
| Multichallenge | — | 20 | the Multichallenge score |
| Polymath | — | 19 | the Polymath score |
| Docvqa | — | 17 | the Docvqa score |
| Graphwalks Bfs 128k | — | 17 | the Graphwalks Bfs 128k score |
| Chartqa | — | 16 | the Chartqa score |
| FrontierCode | — | 16 | the FrontierCode score |
| FrontierCode 1.1 | — | 16 | the FrontierCode 1.1 score |
| Frontierswe | — | 16 | the Frontierswe score |
| Nl2repo | — | 16 | the Nl2repo score |
| Omnidocbench 1 5 | — | 16 | the Omnidocbench 1 5 score |
| Wmt24 | — | 16 | the Wmt24 score |
| Mbpp | — | 15 | the Mbpp score |
| Mmstar | — | 15 | the Mmstar score |
| Odinw | — | 15 | the Odinw score |
| Charxiv D | — | 13 | the Charxiv D score |
| Codeforces | — | 13 | the Codeforces score |
| Graphwalks Parents 128k | — | 13 | the Graphwalks Parents 128k score |
| Cybergym | — | 12 | the Cybergym score |
| Blink | — | 11 | the Blink score |
| Hmmt Feb 26 | — | 11 | the Hmmt Feb 26 score |
| Simplevqa | — | 11 | the Simplevqa score |
| Global Piqa | — | 10 | the Global Piqa score |
| Infovqatest | — | 10 | the Infovqatest score |
| Ocrbench V2 En | — | 10 | the Ocrbench V2 En score |
| Widesearch | — | 10 | the Widesearch score |
| Aider Polyglot Edit | — | 9 | the Aider Polyglot Edit score |
| Babyvision | — | 9 | the Babyvision score |
| Charadessta | — | 9 | the Charadessta score |
| Collie | — | 9 | the Collie score |
| Docvqatest | — | 9 | the Docvqatest score |
| Hiddenmath | — | 9 | the Hiddenmath score |
| Mlvu | — | 9 | the Mlvu score |
| Ocrbench V2 Zh | — | 9 | the Ocrbench V2 Zh score |
| Truthfulqa | — | 9 | the Truthfulqa score |
| Maxife | — | 8 | the Maxife score |
| Mcp Mark | — | 8 | the Mcp Mark score |
| Mlvu M | — | 8 | the Mlvu M score |
| MMMU-Pro (With Tools) | — | 8 | the MMMU-Pro (With Tools) score |
| Arena Agent Elo | — | 7 | the Arena Agent Elo score |
| Artificial Analysis Index | — | 7 | the Artificial Analysis Index score |
| Csimpleqa | — | 7 | the Csimpleqa score |
| Deepsearchqa | — | 7 | the Deepsearchqa score |
| Egoschema | — | 7 | the Egoschema score |
| Natural2code | — | 7 | the Natural2code score |
| Openai Mrcr 2 Needle 128k | — | 7 | the Openai Mrcr 2 Needle 128k score |
| Refcoco Avg | — | 7 | the Refcoco Avg score |
| Textvqa | — | 7 | the Textvqa score |
| V Star | — | 7 | the V Star score |
| Videomme W O Sub | — | 7 | the Videomme W O Sub score |
| Videomme W Sub | — | 7 | the Videomme W Sub score |
| Zebralogic | — | 7 | the Zebralogic score |
| Bird Sql Dev | — | 6 | the Bird Sql Dev score |
| Dynamath | — | 6 | the Dynamath score |
| GRIND | — | 6 | the GRIND score |
| Tau3 Banking | — | 6 | the Tau3 Banking score |
| Androidworld Sr | — | 5 | the Androidworld Sr score |
| Browsecomp Long 128k | — | 5 | the Browsecomp Long 128k score |
| Medxpertqa | — | 5 | the Medxpertqa score |
| Piqa | — | 5 | the Piqa score |
| Swe Lancer Ic Diamond Subset | — | 5 | the Swe Lancer Ic Diamond Subset score |
| Zerobench Sub | — | 5 | the Zerobench Sub score |
| Aa Index | — | 4 | the Aa Index score |
| Boolq | — | 4 | the Boolq score |
| Eclektic | — | 4 | the Eclektic score |
| Exploitgym | — | 4 | the Exploitgym score |
| Fleurs | — | 4 | the Fleurs score |
| Graphwalks Parents (0K-128K) | — | 4 | the Graphwalks Parents (0K-128K) score |
| Hypersim | — | 4 | the Hypersim score |
| Infovqa | — | 4 | the Infovqa score |
| Lingoqa | — | 4 | the Lingoqa score |
| MMMU-Pro (No Tools) | — | 4 | the MMMU-Pro (No Tools) score |
| Mmmu Validation | — | 4 | the Mmmu Validation score |
| Mmvu | — | 4 | the Mmvu score |
| Nexus | — | 4 | the Nexus score |
| OmniDocBench NED | — | 4 | the OmniDocBench NED score |
| Openai Mrcr 2 Needle 1m | — | 4 | the Openai Mrcr 2 Needle 1m score |
| Openbookqa | — | 4 | the Openbookqa score |
| ScienceQA | — | 4 | the ScienceQA score |
| Squality | — | 4 | the Squality score |
| Sunrgbd | — | 4 | the Sunrgbd score |
| Vlmsareblind | — | 4 | the Vlmsareblind score |
| Wmt23 | — | 4 | the Wmt23 score |
| Worldvqa | — | 4 | the Worldvqa score |
| Aider | — | 3 | the Aider score |
| Amc 2022 23 | — | 3 | the Amc 2022 23 score |
| Benchcad With Python Tool | — | 3 | the Benchcad With Python Tool score |
| Capture The Flag Challenges | — | 3 | the Capture The Flag Challenges score |
| Claw Eval Mm | — | 3 | the Claw Eval Mm score |
| Corpusqa 1m | — | 3 | the Corpusqa 1m score |
| Crag | — | 3 | the Crag score |
| Cybersecurity Ctfs | — | 3 | the Cybersecurity Ctfs score |
| Figqa | — | 3 | the Figqa score |
| Finqa | — | 3 | the Finqa score |
| Fullstackbench En | — | 3 | the Fullstackbench En score |
| Fullstackbench Zh | — | 3 | the Fullstackbench Zh score |
| Graphwalks Bfs 1m | — | 3 | the Graphwalks Bfs 1m score |
| Humanitys Last Exam With Tools Text Only | — | 3 | the Humanitys Last Exam With Tools Text Only score |
| Kernelgen 1p | — | 3 | the Kernelgen 1p score |
| Management Consulting Tasks | — | 3 | the Management Consulting Tasks score |
| Math Cot | — | 3 | the Math Cot score |
| Mm Mind2web | — | 3 | the Mm Mind2web score |
| Mrcr V2 8 Needle 512k 1m | — | 3 | the Mrcr V2 8 Needle 512k 1m score |
| Multilingual Mgsm Cot | — | 3 | the Multilingual Mgsm Cot score |
| Multipl E Humaneval | — | 3 | the Multipl E Humaneval score |
| Multipl E Mbpp | — | 3 | the Multipl E Mbpp score |
| Nanogpt | — | 3 | the Nanogpt score |
| Nuscene | — | 3 | the Nuscene score |
| OfficeQA | — | 3 | the OfficeQA score |
| Omniscience | — | 3 | the Omniscience score |
| Openai Connectors | — | 3 | the Openai Connectors score |
| Openai Search Function Calling | — | 3 | the Openai Search Function Calling score |
| Pmc Vqa | — | 3 | the Pmc Vqa score |
| Rsi Index | — | 3 | the Rsi Index score |
| Ruler | — | 3 | the Ruler score |
| Slakevqa | — | 3 | the Slakevqa score |
| Social Iqa | — | 3 | the Social Iqa score |
| Translation En Set1 Comet22 | — | 3 | the Translation En Set1 Comet22 score |
| Translation En Set1 Spbleu | — | 3 | the Translation En Set1 Spbleu score |
| Translation Set1 En Comet22 | — | 3 | the Translation Set1 En Comet22 score |
| Translation Set1 En Spbleu | — | 3 | the Translation Set1 En Spbleu score |
| Usamo 2026 | — | 3 | the Usamo 2026 score |
| Vision2web | — | 3 | the Vision2web score |
| Vqav2 | — | 3 | the Vqav2 score |
| Vqav2 Val | — | 3 | the Vqav2 Val score |
| Xstest | — | 3 | the Xstest score |
| Alpacaeval 2 0 | — | 2 | the Alpacaeval 2 0 score |
| Apex | — | 2 | the Apex score |
| Arxivmath | — | 2 | the Arxivmath score |
| Autologi | — | 2 | the Autologi score |
| Beyond Aime | — | 2 | the Beyond Aime score |
| Bfcl V2 | — | 2 | the Bfcl V2 score |
| Bfcl V3 Multiturn | — | 2 | the Bfcl V3 Multiturn score |
| Browsecomp Long 256k | — | 2 | the Browsecomp Long 256k score |
| Charxiv Reasoning No Tools | — | 2 | the Charxiv Reasoning No Tools score |
| Charxiv Reasoning With Tools | — | 2 | the Charxiv Reasoning With Tools score |
| Cluewsc | — | 2 | the Cluewsc score |
| Covost2 | — | 2 | the Covost2 score |
| Design2code | — | 2 | the Design2code score |
| Dsbench Fullstack | — | 2 | the Dsbench Fullstack score |
| Dsbench Hard | — | 2 | the Dsbench Hard score |
| Factscore | — | 2 | the Factscore score |
| Frames | — | 2 | the Frames score |
| Frontiercode Diamond Xhigh | — | 2 | the Frontiercode Diamond Xhigh score |
| Frontierscience Olympiad | — | 2 | the Frontierscience Olympiad score |
| Frontierscience Research | — | 2 | the Frontierscience Research score |
| Functionalmath | — | 2 | the Functionalmath score |
| Gdm Mrcr V2 8needle 128k Average | — | 2 | the Gdm Mrcr V2 8needle 128k Average score |
| Gdm Mrcr V2 8needle 1m Pointwise | — | 2 | the Gdm Mrcr V2 8needle 1m Pointwise score |
| Gdp Pdf No Tools | — | 2 | the Gdp Pdf No Tools score |
| Graphwalks Parents (256K-1M) | — | 2 | the Graphwalks Parents (256K-1M) score |
| Groundui 1k | — | 2 | the Groundui 1k score |
| Gsm 8k Cot | — | 2 | the Gsm 8k Cot score |
| Horizonmath | — | 2 | the Horizonmath score |
| Humanitys Last Exam No Tools Text Only | — | 2 | the Humanitys Last Exam No Tools Text Only score |
| If | — | 2 | the If score |
| Mmsearch Plus | — | 2 | the Mmsearch Plus score |
| Mobileworld | — | 2 | the Mobileworld score |
| Multilf | — | 2 | the Multilf score |
| Musr | — | 2 | the Musr score |
| OpenAI MRCR v2 8-needle (128K-256K) | — | 2 | the OpenAI MRCR v2 8-needle (128K-256K) score |
| OpenAI MRCR v2 8-needle (128K-512K) | — | 2 | the OpenAI MRCR v2 8-needle (128K-512K) score |
| OpenAI MRCR v2 8-needle (256K-1M) | — | 2 | the OpenAI MRCR v2 8-needle (256K-1M) score |
| OpenAI MRCR v2 8-needle (32K-128K) | — | 2 | the OpenAI MRCR v2 8-needle (32K-128K) score |
| OpenAI MRCR v2 8-needle (4K-8K) | — | 2 | the OpenAI MRCR v2 8-needle (4K-8K) score |
| OpenAI MRCR v2 8-needle (512K-1M) | — | 2 | the OpenAI MRCR v2 8-needle (512K-1M) score |
| OpenAI MRCR v2 8-needle (64K-128K) | — | 2 | the OpenAI MRCR v2 8-needle (64K-128K) score |
| Perceptionbench | — | 2 | the Perceptionbench score |
| Physicsfinals | — | 2 | the Physicsfinals score |
| Polymath En | — | 2 | the Polymath En score |
| Qwen Svg | — | 2 | the Qwen Svg score |
| Superchem | — | 2 | the Superchem score |
| Swe Bench Verified Agentless | — | 2 | the Swe Bench Verified Agentless score |
| Tydiqa | — | 2 | the Tydiqa score |
| Usamo25 | — | 2 | the Usamo25 score |
| Vatex | — | 2 | the Vatex score |
| Videoholmes | — | 2 | the Videoholmes score |
| Visfactor | — | 2 | the Visfactor score |
| Visulogic | — | 2 | the Visulogic score |
| Aa Briefcase Elo | — | 1 | the Aa Briefcase Elo score |
| Aa Omniscience Index | — | 1 | the Aa Omniscience Index score |
| Activitynet | — | 1 | the Activitynet score |
| Advanced Cybersecurity Completion Rate | — | 1 | the Advanced Cybersecurity Completion Rate score |
| Aethercode | — | 1 | the Aethercode score |
| Ai2 Reasoning Challenge Arc | — | 1 | the Ai2 Reasoning Challenge Arc score |
| Aime | — | 1 | the Aime score |
| Aitz Em | — | 1 | the Aitz Em score |
| Android Control High Em | — | 1 | the Android Control High Em score |
| Android Control Low Em | — | 1 | the Android Control Low Em score |
| Arc | — | 1 | the Arc score |
| Arcagi2 | — | 1 | the Arcagi2 score |
| Arkitscenes | — | 1 | the Arkitscenes score |
| Attaq | — | 1 | the Attaq score |
| Babyvision With Python | — | 1 | the Babyvision With Python score |
| Bc Vl | — | 1 | the Bc Vl score |
| Beam 128k | — | 1 | the Beam 128k score |
| Bigcodebench Full | — | 1 | the Bigcodebench Full score |
| Biomysterybench Hard | — | 1 | the Biomysterybench Hard score |
| Biomysterybench Human Solved | — | 1 | the Biomysterybench Human Solved score |
| Cbnsl | — | 1 | the Cbnsl score |
| Cc Bench V2 Backend | — | 1 | the Cc Bench V2 Backend score |
| Cc Bench V2 Frontend | — | 1 | the Cc Bench V2 Frontend score |
| Cc Bench V2 Repo | — | 1 | the Cc Bench V2 Repo score |
| Chartmuseum | — | 1 | the Chartmuseum score |
| Chartqapro | — | 1 | the Chartqapro score |
| Charxiv Rq | — | 1 | the Charxiv Rq score |
| Charxiv Rq With Python | — | 1 | the Charxiv Rq With Python score |
| Ci Memories Coverage | — | 1 | the Ci Memories Coverage score |
| Ci Memories Violation | — | 1 | the Ci Memories Violation score |
| Cloningscenarios | — | 1 | the Cloningscenarios score |
| Codegolf V2 2 | — | 1 | the Codegolf V2 2 score |
| Cohere Agentic Question Answering | — | 1 | the Cohere Agentic Question Answering score |
| Cohere Data Analysis | — | 1 | the Cohere Data Analysis score |
| Cohere Memory Usage Quality | — | 1 | the Cohere Memory Usage Quality score |
| Commonsenseqa | — | 1 | the Commonsenseqa score |
| Contphy | — | 1 | the Contphy score |
| Countqa | — | 1 | the Countqa score |
| Creativework | — | 1 | the Creativework score |
| Crossvid | — | 1 | the Crossvid score |
| Crux O | — | 1 | the Crux O score |
| Cursorbench 3 2 | — | 1 | the Cursorbench 3 2 score |
| Dailyomni | — | 1 | the Dailyomni score |
| Deepsearchqa F1 | — | 1 | the Deepsearchqa F1 score |
| Deepswe 1 0 | — | 1 | the Deepswe 1 0 score |
| Deepswe 1 0 Pass At 1 | — | 1 | the Deepswe 1 0 Pass At 1 score |
| Deepswe 1 1 Mini Swe Agent | — | 1 | the Deepswe 1 1 Mini Swe Agent score |
| Doubao Multi Turn Bench | — | 1 | the Doubao Multi Turn Bench score |
| Draco | — | 1 | the Draco score |
| Ds Arena Code | — | 1 | the Ds Arena Code score |
| Ds Fim Eval | — | 1 | the Ds Fim Eval score |
| Dude | — | 1 | the Dude score |
| Emma | — | 1 | the Emma score |
| Exploitbench Cap Percent | — | 1 | the Exploitbench Cap Percent score |
| Finsearchcomp T2 T3 | — | 1 | the Finsearchcomp T2 T3 score |
| Finsearchcomp T3 | — | 1 | the Finsearchcomp T3 score |
| Flame Vlm Code | — | 1 | the Flame Vlm Code score |
| French Mmlu | — | 1 | the French Mmlu score |
| Frontier Swe Impl | — | 1 | the Frontier Swe Impl score |
| Frontiercode Main | — | 1 | the Frontiercode Main score |
| Frontiercs | — | 1 | the Frontiercs score |
| Gaia2 | — | 1 | the Gaia2 score |
| Gameworld | — | 1 | the Gameworld score |
| Gdp Pdf Mean Criteria No Tools | — | 1 | the Gdp Pdf Mean Criteria No Tools score |
| Gdp Pdf Mean Criteria With Tools | — | 1 | the Gdp Pdf Mean Criteria With Tools score |
| Gdp Pdf Strict Pass | — | 1 | the Gdp Pdf Strict Pass score |
| Govreport | — | 1 | the Govreport score |
| Gpqa Biology | — | 1 | the Gpqa Biology score |
| Gpqa Chemistry | — | 1 | the Gpqa Chemistry score |
| Gpqa Physics | — | 1 | the Gpqa Physics score |
| Harvey Lab Aa | — | 1 | the Harvey Lab Aa score |
| Hipho | — | 1 | the Hipho score |
| Hle Full With Tools | — | 1 | the Hle Full With Tools score |
| Hle Verified | — | 1 | the Hle Verified score |
| Humaneval Er | — | 1 | the Humaneval Er score |
| Image2floorplan | — | 1 | the Image2floorplan score |
| Imagemining | — | 1 | the Imagemining score |
| Infinitebench En Qa | — | 1 | the Infinitebench En Qa score |
| Infographicsqa | — | 1 | the Infographicsqa score |
| Intergps | — | 1 | the Intergps score |
| Kina | — | 1 | the Kina score |
| Livecodebench 01 09 | — | 1 | the Livecodebench 01 09 score |
| Livesports 3k | — | 1 | the Livesports 3k score |
| Loca Bench 256k | — | 1 | the Loca Bench 256k score |
| Longfact Concepts | — | 1 | the Longfact Concepts score |
| Longfact Objects | — | 1 | the Longfact Objects score |
| Lsat | — | 1 | the Lsat score |
| Mask | — | 1 | the Mask score |
| Mathverse | — | 1 | the Mathverse score |
| Mathverse Mini | — | 1 | the Mathverse Mini score |
| Mathvision With Python | — | 1 | the Mathvision With Python score |
| Mbpp Base Version | — | 1 | the Mbpp Base Version score |
| Mbpp Evalplus Base | — | 1 | the Mbpp Evalplus Base score |
| Mbpp Pass 1 | — | 1 | the Mbpp Pass 1 score |
| Mbpp Plus | — | 1 | the Mbpp Plus score |
| Medxpertqa Mm | — | 1 | the Medxpertqa Mm score |
| Mega Mlqa | — | 1 | the Mega Mlqa score |
| Mega Tydi Qa | — | 1 | the Mega Tydi Qa score |
| Mega Udpos | — | 1 | the Mega Udpos score |
| Mega Xcopa | — | 1 | the Mega Xcopa score |
| Mega Xstorycloze | — | 1 | the Mega Xstorycloze score |
| Mewc | — | 1 | the Mewc score |
| Minerva | — | 1 | the Minerva score |
| Mm If Eval | — | 1 | the Mm If Eval score |
| Mmau | — | 1 | the Mmau score |
| Mmbc | — | 1 | the Mmbc score |
| Mmlongbench 128k | — | 1 | the Mmlongbench 128k score |
| Mmlu French | — | 1 | the Mmlu French score |
| Mmmu Pro With Python | — | 1 | the Mmmu Pro With Python score |
| Mmsearch | — | 1 | the Mmsearch score |
| Mmvet | — | 1 | the Mmvet score |
| Mobileminiwob Sr | — | 1 | the Mobileminiwob Sr score |
| Mrcr 128k 8 Needle | — | 1 | the Mrcr 128k 8 Needle score |
| Mrcr 1m Pointwise | — | 1 | the Mrcr 1m Pointwise score |
| Msqa | — | 1 | the Msqa score |
| Mt Aime 2025 | — | 1 | the Mt Aime 2025 score |
| Next.js Evals + AGENTS.md | — | 1 | the Next.js Evals + AGENTS.md score |
| Objectron | — | 1 | the Objectron score |
| Officeqa Pro Vision | — | 1 | the Officeqa Pro Vision score |
| Ojbench Cpp | — | 1 | the Ojbench Cpp score |
| Omniscience Non Hallucination Rate | — | 1 | the Omniscience Non Hallucination Rate score |
| Open Rewrite | — | 1 | the Open Rewrite score |
| Openai Mrcr 2 Needle 256k | — | 1 | the Openai Mrcr 2 Needle 256k score |
| Openrca | — | 1 | the Openrca score |
| Osworld Extended | — | 1 | the Osworld Extended score |
| Osworld Screenshot Only | — | 1 | the Osworld Screenshot Only score |
| Perceptiontest | — | 1 | the Perceptiontest score |
| Phibench | — | 1 | the Phibench score |
| Pope | — | 1 | the Pope score |
| Popqa | — | 1 | the Popqa score |
| Prbench Legal | — | 1 | the Prbench Legal score |
| Protocolqa | — | 1 | the Protocolqa score |
| Qasper | — | 1 | the Qasper score |
| Qmsum | — | 1 | the Qmsum score |
| Qvhighlights | — | 1 | the Qvhighlights score |
| Qwen Qoder Bench | — | 1 | the Qwen Qoder Bench score |
| Qwen React Bench | — | 1 | the Qwen React Bench score |
| Realkie Fcc | — | 1 | the Realkie Fcc score |
| Recreationbench | — | 1 | the Recreationbench score |
| Repo Env | — | 1 | the Repo Env score |
| Repoqa | — | 1 | the Repoqa score |
| Robospatialhome | — | 1 | the Robospatialhome score |
| Sat Math | — | 1 | the Sat Math score |
| Scienceqa Visual | — | 1 | the Scienceqa Visual score |
| Seedclawbench | — | 1 | the Seedclawbench score |
| Sifo | — | 1 | the Sifo score |
| Sifo Multiturn | — | 1 | the Sifo Multiturn score |
| Siren Agentdojo Attack Success | — | 1 | the Siren Agentdojo Attack Success score |
| Siren Agentdojo Utility | — | 1 | the Siren Agentdojo Utility score |
| Spider | — | 1 | the Spider score |
| Summscreenfd | — | 1 | the Summscreenfd score |
| Superglue | — | 1 | the Superglue score |
| Surds | — | 1 | the Surds score |
| Swe Bench Pro Resolve Rate | — | 1 | the Swe Bench Pro Resolve Rate score |
| Swe Bench Verified Multiple Attempts | — | 1 | the Swe Bench Verified Multiple Attempts score |
| Swe Marathon Resolution Rate | — | 1 | the Swe Marathon Resolution Rate score |
| Tau3 Airline | — | 1 | the Tau3 Airline score |
| Tau3 Retail | — | 1 | the Tau3 Retail score |
| Tau3 Telecom | — | 1 | the Tau3 Telecom score |
| Tempcompass | — | 1 | the Tempcompass score |
| Terminus | — | 1 | the Terminus score |
| Theoremqa | — | 1 | the Theoremqa score |
| Tldr9 Test | — | 1 | the Tldr9 Test score |
| Tomato | — | 1 | the Tomato score |
| Trae Code Gen | — | 1 | the Trae Code Gen score |
| Trae Error Fix | — | 1 | the Trae Error Fix score |
| Uniform Bar Exam | — | 1 | the Uniform Bar Exam score |
| Vct | — | 1 | the Vct score |
| Vibe | — | 1 | the Vibe score |
| Vibe Android | — | 1 | the Vibe Android score |
| Vibe Backend | — | 1 | the Vibe Backend score |
| Vibe Ios | — | 1 | the Vibe Ios score |
| Vibe Simulation | — | 1 | the Vibe Simulation score |
| Vibe Web | — | 1 | the Vibe Web score |
| Video Mme Long No Subtitles | — | 1 | the Video Mme Long No Subtitles score |
| Videosimpleqa | — | 1 | the Videosimpleqa score |
| Vlmsarebiased | — | 1 | the Vlmsarebiased score |
| Voicebench Avg | — | 1 | the Voicebench Avg score |
| Vqav2 Test | — | 1 | the Vqav2 Test score |
| We Math | — | 1 | the We Math score |
| Webvoyager | — | 1 | the Webvoyager score |
| Wmdp | — | 1 | the Wmdp score |
| Worldvqa Forceanswer | — | 1 | the Worldvqa Forceanswer score |
| Zerobench Main Pass At 5 | — | 1 | the Zerobench Main Pass At 5 score |
| Zerobench Main With Python Pass At 5 | — | 1 | the Zerobench Main With Python Pass At 5 score |
