GPQA Diamond
GPQA is a set of multiple-choice science questions written and checked by people with or working toward a PhD in the field. The Diamond subset keeps the 198 questions that both expert checkers got right and that most skilled non-experts got wrong, even with web access. Each question has four options, so random guessing scores 25%. It measures careful scientific reasoning in text. It says nothing about coding, tool use, speed, or how a model behaves once quantized.
Sources: GPQA paper (arXiv 2311.12022) · OpenRouter GPQA Diamond benchmark
against memory
OpenRouter evals · data from 2026-10-10
Not beaten on both memory and score● memory stated by the publisher◇ memory estimated
| # | Model | Configuration | Min memory | GPQA Diamond | Frontier |
|---|---|---|---|---|---|
| 1 | Qwen3.8 27B | 27B Original weights | 54.3 GB ◇ | 81.9% | On frontier |
| 2 | Nemotron 3.5 Lightning | 30B-A3B Original weights BF16 | 62.3 GB ◇ | 69.8% | — |
| 3 | DeepSeek V4 Flash Vision Exp | Original weights | 158 GB ◇ | 88.0% | On frontier |
| 4 | Ling 3.0 Flash VL | Original weights | 238 GB ◇ | 84.8% | — |
| 5 | GLM 5.3 Flash | Original weights | 306 GB ◇ | 85.0% | — |
| 6 | Qwen3.8 Flash-Next | Original weights | 337 GB ◇ | 88.6% | On frontier |
| 7 | Step 3.7 Flash | Original weights | 387 GB ◇ | 76.6% | — |
| 8 | DeepSeek V4.1 Flash | Original weights | 476 GB ◇ | 88.8% | On frontier |
| 9 | Kimi K2.7 Code | Original weights | 568 GB ◇ | 86.0% | — |
| 10 | GLM 5.3 | Original weights | 734 GB ◇ | 85.2% | — |
| 11 | MiniMax M3 | Original weights | 797 GB ◇ | 90.5% | On frontier |
| 12 | DeepSeek V4 Pro | Original weights | 807 GB ◇ | 90.1% | — |
| 13 | Nemotron 3 Ultra | 550B-A55B Original weights BF16 | 1045 GB ◇ | 81.2% | — |
| 14 | Hy4 preview | Original weights | 1455 GB ◇ | 89.7% | — |
| 15 | Kimi K3 | Original weights | 1475 GB ◇ | 92.0% | On frontier |
| 16 | Qwen3.8 2.4T-A95B | 2.4T-A95B Original weights | 4560 GB ◇ | 86.9% | — |
Scores were measured on hosted endpoints, so each model is plotted once, at the memory need of its original weights. A quantized file needs less memory and may score differently. Lines connect the frontier points and do not predict results in between. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai
Open-weight models
| Model | Endpoint | Score | Tasks run | Original weights |
|---|---|---|---|---|
| Kimi K3 | moonshotai/kimi-k3-20260715 | 92.0% ± 1.2 | 1,974 | 1475 GB |
| GLM 5.3 Flash | z-ai/glm-5.3-flash | 90.9% ± 0.0 | 198 | 306 GB |
| MiniMax M3 | minimax/minimax-m3-20260531 | 90.5% ± 1.9 | 2,373 | 797 GB |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro-20260813 | 90.1% ± 1.7 | 1,779 | 807 GB |
| Hy4 preview | tencent/hy4-preview-20260827 | 89.7% ± 0.0 | 198 | 1455 GB |
| DeepSeek V4.1 Flash | deepseek/deepseek-v4.1-flash-20260910 | 88.8% ± 0.9 | 792 | 476 GB |
| Qwen3.8 Flash-Next | qwen/qwen3.8-flash-20260826 | 88.6% ± 0.0 | 198 | 337 GB |
| DeepSeek V4 Flash Vision Exp | deepseek/deepseek-v4-flash-vision-exp-20260821 | 88.0% ± 1.1 | 1,188 | 158 GB |
| DeepSeek V4.1 Flash | deepseek/deepseek-v4.1-flash | 87.9% ± 0.0 | 198 | 476 GB |
| Qwen3.8 2.4T-A95B | qwen/qwen3.8-2.4t-a95b-20260812 | 86.9% ± 1.2 | 1,386 | 4560 GB |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro-20260423 | 86.4% ± 1.6 | 1,584 | 807 GB |
| Kimi K2.7 Code | moonshotai/kimi-k2.7-code-20260612 | 86.0% ± 1.8 | 1,980 | 568 GB |
| GLM 5.3 | z-ai/glm-5.3-20260816 | 85.2% ± 1.4 | 1,386 | 734 GB |
| GLM 5.3 Flash | z-ai/glm-5.3-flash-20260826 | 85.0% ± 2.3 | 1,386 | 306 GB |
| Ling 3.0 Flash VL | inclusionai/ling-3.0-flash-vl-20260910 | 84.8% ± 0.0 | 198 | 238 GB |
| GLM 5.3 | z-ai/glm-5.3 | 82.3% ± 0.0 | 198 | 734 GB |
| Qwen3.8 27B | qwen/qwen3.8-27b-20260814 | 81.9% ± 1.5 | 1,584 | 54.3 GB |
| Nemotron 3 Ultra | nvidia/nemotron-3-ultra-550b-a55b-20260604 | 81.2% ± 0.0 | 172 | 1045 GB |
| Step 3.7 Flash | stepfun/step-3.7-flash-20260528 | 76.6% ± 0.0 | 198 | 387 GB |
| Nemotron 3.5 Lightning | nvidia/nemotron-3.5-lightning-20260807 | 69.8% ± 2.8 | 1,530 | 62.3 GB |
Measured by OpenRouter evals on hosted endpoints, not by this site. A quantized file you run yourself may score differently. One model can have several endpoints. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai