GPQA Diamond

GPQA is a set of multiple-choice science questions written and checked by people with or working toward a PhD in the field. The Diamond subset keeps the 198 questions that both expert checkers got right and that most skilled non-experts got wrong, even with web access. Each question has four options, so random guessing scores 25%. It measures careful scientific reasoning in text. It says nothing about coding, tool use, speed, or how a model behaves once quantized.

Sources: GPQA paper (arXiv 2311.12022) · OpenRouter GPQA Diamond benchmark

against memory

OpenRouter evals · data from 2026-10-10

16 of 16 models
6570758085909532641282565121024204840968192GPQA Diamond (%) ↑Minimum memory (GB) · log scale12345678910111213141516

Not beaten on both memory and score● memory stated by the publisher◇ memory estimated

#ModelConfigurationMin memoryGPQA DiamondFrontier
1Qwen3.8 27B27B Original weights54.3 GB ◇81.9%On frontier
2Nemotron 3.5 Lightning30B-A3B Original weights BF1662.3 GB ◇69.8%—
3DeepSeek V4 Flash Vision ExpOriginal weights158 GB ◇88.0%On frontier
4Ling 3.0 Flash VLOriginal weights238 GB ◇84.8%—
5GLM 5.3 FlashOriginal weights306 GB ◇85.0%—
6Qwen3.8 Flash-NextOriginal weights337 GB ◇88.6%On frontier
7Step 3.7 FlashOriginal weights387 GB ◇76.6%—
8DeepSeek V4.1 FlashOriginal weights476 GB ◇88.8%On frontier
9Kimi K2.7 CodeOriginal weights568 GB ◇86.0%—
10GLM 5.3Original weights734 GB ◇85.2%—
11MiniMax M3Original weights797 GB ◇90.5%On frontier
12DeepSeek V4 ProOriginal weights807 GB ◇90.1%—
13Nemotron 3 Ultra550B-A55B Original weights BF161045 GB ◇81.2%—
14Hy4 previewOriginal weights1455 GB ◇89.7%—
15Kimi K3Original weights1475 GB ◇92.0%On frontier
16Qwen3.8 2.4T-A95B2.4T-A95B Original weights4560 GB ◇86.9%—

Scores were measured on hosted endpoints, so each model is plotted once, at the memory need of its original weights. A quantized file needs less memory and may score differently. Lines connect the frontier points and do not predict results in between. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai

Open-weight models

OpenRouter evals · data from 2026-10-10
ModelEndpointScoreTasks runOriginal weights
Kimi K3moonshotai/kimi-k3-2026071592.0% ± 1.21,9741475 GB
GLM 5.3 Flashz-ai/glm-5.3-flash90.9% ± 0.0198306 GB
MiniMax M3minimax/minimax-m3-2026053190.5% ± 1.92,373797 GB
DeepSeek V4 Prodeepseek/deepseek-v4-pro-2026081390.1% ± 1.71,779807 GB
Hy4 previewtencent/hy4-preview-2026082789.7% ± 0.01981455 GB
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash-2026091088.8% ± 0.9792476 GB
Qwen3.8 Flash-Nextqwen/qwen3.8-flash-2026082688.6% ± 0.0198337 GB
DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp-2026082188.0% ± 1.11,188158 GB
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash87.9% ± 0.0198476 GB
Qwen3.8 2.4T-A95Bqwen/qwen3.8-2.4t-a95b-2026081286.9% ± 1.21,3864560 GB
DeepSeek V4 Prodeepseek/deepseek-v4-pro-2026042386.4% ± 1.61,584807 GB
Kimi K2.7 Codemoonshotai/kimi-k2.7-code-2026061286.0% ± 1.81,980568 GB
GLM 5.3z-ai/glm-5.3-2026081685.2% ± 1.41,386734 GB
GLM 5.3 Flashz-ai/glm-5.3-flash-2026082685.0% ± 2.31,386306 GB
Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl-2026091084.8% ± 0.0198238 GB
GLM 5.3z-ai/glm-5.382.3% ± 0.0198734 GB
Qwen3.8 27Bqwen/qwen3.8-27b-2026081481.9% ± 1.51,58454.3 GB
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b-2026060481.2% ± 0.01721045 GB
Step 3.7 Flashstepfun/step-3.7-flash-2026052876.6% ± 0.0198387 GB
Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning-2026080769.8% ± 2.81,53062.3 GB

Measured by OpenRouter evals on hosted endpoints, not by this site. A quantized file you run yourself may score differently. One model can have several endpoints. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai