Who ranks #1 on the Science leaderboard?
As of August 21, 2026, GPT-5.6 Sol by OpenAI ranks #1 on GPQA at 94.6%. API pricing is $5.00/M input and $30.00/M output.
As of August 21, 2026, GPT-5.6 Sol is #1 for research and science at 94.6%. Ranked by GPQA: graduate-level science questions that resist simple search. Best AI for research and science, ranked by GPQA Diamond, HealthBench, SciCode, and FrontierMath. Compare research scores next to API price.
![]() GPT-5.6 Sol OpenAI ยท Proprietary | 94.6% | 82.8% | 56.9% | 32.3% | 44.0% | 85.2% | 83% | 89% | โ | โ | โ | 57% | $35.00 |
![]() Claude Mythos Preview Anthropic ยท Proprietary | 94.6% | โ | โ | โ | โ | โ | โ | โ | โ | 93.2% | โ | โ | $60.00 |
![]() Claude Mythos 5 Anthropic ยท Proprietary | 94.6% | 94.6% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $60.00 |
Gemini 3.1 Pro Google ยท Proprietary | 94.3% | 94.4% | 59% | 17.7% | 59.1% | 76.1% | 80.5% | 36.9% | 96.4% | โ | 16.7% | โ | $17.50 |
![]() Claude Opus 4.7 Anthropic ยท Proprietary | 94.2% | 86.4% | 54.5% | 12% | 54.9% | 83.0% | โ | 43.8% | โ | 91% | 22.9% | โ | $30.00 |
![]() GPT-5.5 OpenAI ยท Proprietary | 93.6% | 77.3% | 56.1% | 27.1% | 49.1% | 86.9% | 83.2% | 35.4% | โ | โ | 35.4% | โ | $35.00 |
![]() Claude Opus 4.8 Anthropic ยท Proprietary | 93.6% | 85.3% | 53.5% | 20.9% | 53.2% | 85.8% | โ | 47.2% | โ | 89.9% | 31.3% | โ | $30.00 |
Kimi K3OSS Moonshot AI ยท Open Source | 93.5% | 93.5% | 58.7% | 23.4% | 48.9% | 88.0% | 81.6% | โ | โ | 91.3% | โ | โ | $18.00 |
![]() GPT-5.2 Pro OpenAI ยท Proprietary | 93.2% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $189.00 |
![]() Grok 4.5 xAI ยท Proprietary | 93% | 93.4% | 54.0% | 15.4% | 43.3% | 86.9% | โ | โ | โ | โ | โ | โ | $8.00 |
![]() GPT-5.6 Terra OpenAI ยท Proprietary | 92.9% | 77.3% | 53.9% | 30% | 43.4% | 82.9% | 80.7% | 84.9% | โ | โ | โ | 57% | $14.00 |
![]() GPT-5.4 OpenAI ยท Proprietary | 92.8% | 89.4% | 56.6% | 23.4% | 41.3% | 77.5% | 81.2% | 47.6% | 96.1% | โ | 27.1% | โ | $17.50 |
Qwen3.8 MaxOSS Qwen ยท Open Source | 92.6% | 92.7% | โ | โ | 40.7% | 85.0% | 82.3% | โ | โ | โ | โ | 60.2% | $8.00 |
Qwen3.7 Max Qwen ยท Proprietary | 92.4% | 90.9% | 53.5% | 11.4% | 38.8% | 79.4% | โ | โ | โ | โ | โ | โ | $5.00 |
![]() GPT-5.2 OpenAI ยท Proprietary | 92.4% | 73.2% | โ | โ | 49.8% | 84.4% | 79.5% | 40.3% | 94.1% | 82.1% | 18.8% | โ | $15.75 |
![]() GPT-5.6 Luna OpenAI ยท Proprietary | 92.3% | 63.6% | 52.5% | 20.6% | 42.4% | 84.4% | 78.4% | 78.6% | โ | โ | โ | 55.8% | $1.40 |
Gemini 3 Pro Google ยท Proprietary | 91.9% | 91.9% | โ | 6.9% | 52.2% | 72.0% | 81% | 37.6% | 96.0% | 81.4% | 18.8% | โ | $14.00 |
![]() Claude Opus 4.6 Anthropic ยท Proprietary | 91.3% | 88.4% | โ | โ | 48.2% | 86.7% | 77.3% | 40.7% | โ | 77.4% | 22.9% | โ | $30.00 |
GLM-5.2OSS Z AI ยท Open Source | 91.2% | 71.2% | 50.5% | 16.7% | 40.8% | 83.5% | โ | โ | โ | โ | โ | โ | $5.80 |
Kimi K2.6OSS Moonshot AI ยท Open Source | 90.5% | 90.8% | 52.2% | 8% | 40.1% | 78.2% | 80.1% | 39.0% | โ | 86.7% | 14.6% | โ | $4.93 |
Qwen3.6 Plus Qwen ยท Proprietary | 90.4% | 88.4% | 40.7% | 2.9% | 36.9% | 77.0% | 78.8% | 26.2% | โ | 81.5% | 8.3% | โ | $3.50 |
Hy3OSS Tencent ยท Open Source | 90.4% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $0.66 |
Gemini 3 Flash Google ยท Proprietary | 90.4% | 89.4% | โ | โ | 55.9% | 69.9% | 81.2% | 35.6% | 95.8% | 80.3% | 4.2% | โ | $3.50 |
Qwen3.7-Plus Qwen ยท Proprietary | 90.3% | 87.9% | 51.3% | 6% | โ | โ | 79% | โ | โ | 85.9% | โ | โ | $1.60 |
DeepSeek-V4-Pro-MaxOSS DeepSeek ยท Open Source | 90.1% | 89.7% | 50% | 12.9% | โ | โ | โ | โ | โ | โ | โ | โ | $5.22 |
![]() Claude Sonnet 4.6 Anthropic ยท Proprietary | 89.9% | 78.8% | 46.8% | 3.1% | โ | โ | 75.6% | 32.4% | 92.1% | โ | 8.3% | โ | $18.00 |
Inkling-SmallOSS Thinking Machines ยท Open Source ยท via Thinking Machines Lab | 89.5% | 88.5% | 48.7% | 8.3% | 37.9% | 84.1% | 74% | โ | โ | 77.4% | โ | โ | $1.50 |
Qwen3.8-27BOSS Qwen ยท Open Source | 89.2% | โ | โ | โ | โ | โ | โ | โ | โ | 90.2% | โ | โ | $3.65 |
Solar Pro 4 Upstage ยท Proprietary | 89% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $1.50 |
Seed 2.0 Pro ByteDance ยท Proprietary | 88.9% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $3.50 |
Qwen3.5-397B-A17BOSS Qwen ยท Open Source | 88.4% | 85.9% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $4.20 |
![]() GPT-5.1 Thinking OpenAI ยท Proprietary | 88.1% | โ | โ | โ | โ | โ | โ | 26.7% | โ | โ | โ | โ | $11.25 |
![]() GPT-5.1 Instant OpenAI ยท Proprietary | 88.1% | โ | โ | โ | โ | โ | โ | 26.7% | โ | โ | โ | โ | $11.25 |
![]() GPT-5.1 High OpenAI ยท Proprietary | 88.1% | โ | 43.3% | 4.9% | โ | โ | โ | โ | โ | โ | โ | โ | $11.25 |
![]() GPT-5.1 OpenAI ยท Proprietary | 88.1% | 66.7% | 43.3% | 4.9% | 52.7% | 88.1% | โ | 26.7% | 96.4% | โ | 4.2% | โ | $11.25 |
![]() GPT-5 Medium OpenAI ยท Proprietary | 88.1% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $11.25 |
DeepSeek-V4-Flash-MaxOSS DeepSeek ยท Open Source | 88.1% | โ | 44.9% | 7.1% | โ | โ | โ | โ | โ | โ | โ | โ | $0.42 |
![]() GPT-5.4 mini OpenAI ยท Proprietary | 88% | 88% | 49.9% | 10% | โ | โ | 76.6% | 28.3% | โ | โ | 2.1% | โ | $5.25 |
Qwen3.6-27BOSS Qwen ยท Open Source | 87.8% | 85.9% | โ | โ | โ | โ | 75.8% | โ | โ | 78.4% | โ | โ | $4.20 |
![]() GPT-5.3 Codex OpenAI ยท Proprietary | 87.7% | 87.7% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $15.75 |
Kimi K2.5OSS Moonshot AI ยท Open Source | 87.6% | โ | 48.7% | 3.1% | โ | โ | 78.5% | โ | โ | 77.5% | โ | โ | $3.68 |
![]() Grok-4 xAI ยท Proprietary | 87.5% | 87.5% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $18.00 |
DeepSeek-V4-Flash-0423OSS DeepSeek ยท Open Source | 87.4% | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | โ | $0.30 |
![]() GPT-5 High OpenAI ยท Proprietary | 87.3% | โ | 42.9% | โ | โ | โ | โ | โ | โ | โ | โ | โ | $11.25 |
Nemotron 3 Ultra (550B A55B)OSS NVIDIA ยท Open Source ยท via OpenRouter | 87% | 86.1% | 44.6% | 3.1% | 38.6% | โ | โ | โ | โ | โ | โ | โ | $2.70 |
![]() Claude Opus 4.5 Anthropic ยท Proprietary | 87% | 85.5% | โ | โ | 45.2% | 83.3% | โ | 20.7% | 93.2% | โ | 4.2% | โ | $30.00 |
Gemini 3.1 Flash-Lite Google ยท Proprietary | 86.9% | 81.8% | 41.9% | 1.1% | 47.6% | 63.9% | 76.8% | โ | โ | 73.2% | โ | โ | $1.75 |
Qwen3.5-122B-A10BOSS Qwen ยท Open Source | 86.6% | โ | โ | โ | โ | โ | 76.9% | โ | โ | 77.2% | โ | โ | $3.60 |
Gemini 2.5 Pro Preview 06-05 Google ยท Proprietary | 86.4% | 84.8% | โ | โ | โ | โ | โ | 10.3% | โ | โ | 2.1% | โ | $11.25 |
GLM-5.1OSS Z AI ยท Open Source | 86.2% | 89.9% | 43.8% | 4.6% | 41.6% | 72.3% | โ | 33.5% | โ | โ | 12.5% | โ | $5.80 |
Next step
Auth, billing, and the API layer are already decided. Fork 8 finished AI apps, or have us build the first version with you.
The index
282 models across 33 providers. Search or jump to a lab โ every model page stays linked here.







As of August 21, 2026, GPT-5.6 Sol by OpenAI is #1 for research and science at 94.6%. Ranked by GPQA: graduate-level science questions that resist simple search. This board also tracks GPQA, GPQA Diamond, SciCode, CritPt. Next on the same board: Claude Mythos 5 and Claude Mythos Preview. Related leaders: Gemini 3.7 Flash on GPQA Diamond at 94.8%; Claude Fable 5 on SciCode at 60.2%. This science leaderboard ranks models by GPQA. Scores come from public evals. Prices are the live API rates in the table above.
Sources: GPQA (GitHub) (https://github.com/idavidrein/gpqa); OpenAI API pricing (https://developers.openai.com/api/docs/pricing); Anthropic Claude API pricing (https://platform.claude.com/docs/en/about-claude/pricing); Gemini API pricing (https://ai.google.dev/gemini-api/docs/pricing)
| Rank | Model | GPQA | Input /M | Output /M |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 94.6% | $5.00 | $30.00 |
| 2 | Claude Mythos 5 | 94.6% | $10.00 | $50.00 |
| 3 | Claude Mythos Preview | 94.6% | $10.00 | $50.00 |
| 4 | Gemini 3.1 Pro | 94.3% | $2.50 | $15.00 |
| 5 | Claude Opus 4.7 | 94.2% | $5.00 | $25.00 |
| 6 | Claude Opus 4.8 | 93.6% | $5.00 | $25.00 |
| 7 | GPT-5.5 | 93.6% | $5.00 | $30.00 |
| 8 | Kimi K3 | 93.5% | $3.00 | $15.00 |
People search best AI for research more than best AI for science. This board uses GPQA Diamond for graduate science, SciCode for scientific programming, HealthBench for medical dialogue, and chart evals for papers with figures. Deep-research products are agents on top of a model. Rank the model here.
HealthBench scores how a model handles health conversations. It is closer to a clinician chat than to GPQA. If your product is medical, sort HealthBench. If your product is physics QA, sort GPQA Diamond.
Rank one eval at a time. All LLM benchmarks.
As of August 21, 2026, GPT-5.6 Sol by OpenAI ranks #1 on GPQA at 94.6%. API pricing is $5.00/M input and $30.00/M output.
As of August 21, 2026, GPT-5.6 Sol by OpenAI is #1 for research and science at 94.6%. Ranked by GPQA: graduate-level science questions that resist simple search. This board also tracks GPQA, GPQA Diamond, SciCode, CritPt. Next on the same board: Claude Mythos 5 and Claude Mythos Preview. Related leaders: Gemini 3.7 Flash on GPQA Diamond at 94.8%; Claude Fable 5 on SciCode at 60.2%.
The current GPQA ranking as of August 21, 2026 is 1. GPT-5.6 Sol at 94.6%; 2. Claude Mythos 5 at 94.6%; 3. Claude Mythos Preview at 94.6%.
Llama 3.2 3B Instruct is the cheapest scored model on this science leaderboard at $0.01/M input and $0.02/M output ($0.03 blended). GPT-5.6 Sol still leads GPQA at 94.6%.
Not automatically. GPT-5.6 Sol leads GPQA, but a cheaper scored model can be the better production choice if the quality gap is small. Use the table to weigh GPQA against input/output price, context window, and related evals.
Scores and API prices on this page are refreshed from published evals and provider rates. The snapshot is labeled August 21, 2026. Treat it as a current index, not a one-off blog post.
GPQA Diamond and FrontierMath rank scientific reasoning. SciCode ranks scientific programming. HealthBench ranks medical dialogue.
HealthBench is OpenAI's medical conversation eval. It scores how models handle clinical questions, not just multiple-choice science.
Yes. This page keeps the science evals next to input and output prices so a lab model can be judged on both accuracy and spend.
Best AI for science on this page means GPQA Diamond, Frontier Science, SciCode, HealthBench, and chart/vision science evals, next to API price. Best AI for research is the higher-volume version of the same question.
Use HealthBench and MedQA as a public signal, then run your own eval. This is not clinical advice and it is not a device. Token price still matters if you screen a lot of papers.