Local LLMs

47 current open-weight models from 22 labs. Latest version of each family only.

against memory

OpenRouter evals, GPQA Diamond · data from 2026-10-10

15 of 15 models
6570758085909532641282565121024204840968192GPQA Diamond (%) ↑Minimum memory (GB) · log scale123456789101112131415

Not beaten on both memory and score● memory stated by the publisher◇ memory estimated

#ModelConfigurationMin memoryGPQA DiamondFrontier
1Qwen3.8 27B27B Original weights54.3 GB ◇81.9%On frontier
2Nemotron 3.5 Lightning30B-A3B Original weights BF1662.3 GB ◇69.8%—
3DeepSeek V4 Flash Vision ExpOriginal weights158 GB ◇88.0%On frontier
4Ling 3.0 Flash VLOriginal weights238 GB ◇84.8%—
5GLM 5.3 FlashOriginal weights306 GB ◇85.0%—
6Qwen3.8 Flash-NextOriginal weights337 GB ◇88.6%On frontier
7Step 3.7 FlashOriginal weights387 GB ◇76.6%—
8DeepSeek V4.1 FlashOriginal weights476 GB ◇88.8%On frontier
9GLM 5.3Original weights734 GB ◇85.2%—
10MiniMax M3Original weights797 GB ◇90.5%On frontier
11DeepSeek V4 ProOriginal weights807 GB ◇90.1%—
12Nemotron 3 Ultra550B-A55B Original weights BF161045 GB ◇81.2%—
13Hy4 previewOriginal weights1455 GB ◇89.7%—
14Kimi K3Original weights1475 GB ◇92.0%On frontier
15Qwen3.8 2.4T-A95B2.4T-A95B Original weights4560 GB ◇86.9%—

Scores were measured on hosted endpoints, so each model is plotted once, at the memory need of its original weights. A quantized file needs less memory and may score differently. Lines connect the frontier points and do not predict results in between. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai

Models

47
ModelLabParamsYour deviceScore
Kimi K3Moonshot AI2.78T1475 GB ◇92.0% GPQAMiniMax M3MiniMax427B797 GB ◇90.5% GPQADeepSeek V4 ProDeepSeek1.6T807 GB ◇90.1% GPQAHy4 previewTencent780B1455 GB ◇89.7% GPQADeepSeek V4.1 FlashDeepSeek763.2B476 GB ◇88.8% GPQAQwen3.8 Flash-NextAlibaba Qwen180B337 GB ◇88.6% GPQADeepSeek V4 Flash Vision ExpPreviewDeepSeek304.6B158 GB ◇88.0% GPQAQwen3.8 2.4T-A95BAlibaba Qwen2.45T4560 GB ◇86.9% GPQAGLM 5.3Z.ai753.3B734 GB ◇85.2% GPQAGLM 5.3 FlashZ.ai321.3B306 GB ◇85.0% GPQALing 3.0 Flash VLAnt Group inclusionAI124.8B238 GB ◇84.8% GPQAQwen3.8 27BAlibaba Qwen27.8B54.3 GB ◇81.9% GPQANemotron 3 UltraNVIDIA560.5B1045 GB ◇81.2% GPQAStep 3.7 FlashStepFun201.4B82.4 GB ◇76.6% GPQANemotron 3.5 LightningNVIDIA31.6B62.3 GB ◇69.8% GPQALensVLM 9BApple9.41B19.1 GB ◇MiMo V2.6Xiaomi9.41B19.1 GB ◇FrogNano 4BMicrosoft4.66B10.2 GB ◇AstaBrief 8B SFTAi2—16.9 GB ◇LLaDA 2.2Ant Group inclusionAI16.3B31.1 GB ◇tiny-aya ThinkerPreviewCohere3.35B6.8 GB ◇Granite 4.2IBM3.66B2.5 GB ◇Ling 3.0Ant Group inclusionAI7.89B6.5 GB ◇North Micro VisionPreviewCohere2.48B6.1 GB ◇LongCat Flash Lite SparseMeituan69.1B129 GB ◇K-EXAONE 2.0LG AI Research749.4B1399 GB ◇Mage-VLMicrosoft4.74B11.5 GB ◇LongCat 2.0Meituan1.78T3308 GB ◇Granite SwashIBM2.14B4.9 GB ◇Nemotron Labs 3 PuzzlePreviewNVIDIA75.4B146 GB ◇UltraX 0.6B PreviewOpenBMB—3.9 GB ◇DiffusionGemmaGoogle25.8B50.5 GB ◇Gemma 4 12BGoogle12B25.8 GB ◇MiniCPM5OpenBMB1.08B1.3 GB ◇BitCPM-CANNOpenBMB—0.9 GB ◇Ring 2.6 1TAnt Group inclusionAI1.03T991 GB ◇Command A Plus 05-2026PreviewCohere218.8B409 GB ◇Nemotron Labs Diffusion VLMPreviewNVIDIA8.92B18.2 GB ◇Granite Switch 4.1 PreviewIBM4.15B8.9 GB ◇Mistral Medium 3.5Mistral AI127.7B252 GB ◇nanowhale 100MHugging Face110M1.0 GB ◇LLaDA 2.0 UniAnt Group inclusionAI16.3B56.6 GB ◇Nemotron Labs DiffusionPreviewNVIDIA13.5B27.1 GB ◇rnj 1.5Essential AI8.31B32.5 GB ◇Nemotron 3 Nano OmniNVIDIA33B62.5 GB ◇BARAi27.3B18.1 GB ◇MiniCPM-V 4.6OpenBMB1.3B1.4 GB ◇

Models with a GPQA Diamond score come first, highest score first. The rest are newest first.

Min memory is the smallest listed option. ● stated by the publisher · ◇ estimate from file size, an 8K-token context and 0.5 GB for the runtime. — means no figure is available.