Local LLMs that fit in 12 GB
12 of 47 current models have a listed weight option that needs at most 85% of 12 GB.
against memory
OpenRouter evals, GPQA Diamond · data from 2026-10-10
No model with a score fits this memory budget.
Scores were measured on hosted endpoints, so each model is plotted once, at the memory need of its original weights. A quantized file needs less memory and may score differently. Lines connect the frontier points and do not predict results in between. Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). openrouter.ai
Models
Models with a GPQA Diamond score come first, highest score first. The rest are newest first.
An option is listed when its memory need is at most 85% of 12 GB. The figure after each option is an estimate from file size, an 8K-token context and 0.5 GB for the runtime, unless the publisher states one.