Method
Where the list comes from, how the numbers are made, and what has been tested.
What is listed
The list covers open-weight models whose weights can be downloaded from the publisher's own Hugging Face organization. Each model family appears once, at its latest version. When a lab ships a new version, the older one leaves the list and its address redirects to the new page. Sizes of the same generation share one page. A preview release is listed and marked Preview. A model that has only been announced is not listed.
Dates are the publisher's release date when one is given, otherwise the day the repository was created; each page says which. Parameter counts are the total in the weight files of the listed repository. For a mixture-of-experts model that is every expert, not the active part. Licenses are the identifiers from the repository. Read the license text before you build on a model.
The list was last checked on 2026-10-11: 154 models from 44 labs.
Weight options
A weight option is one size, one format and one quantization level, for example 8B, GGUF, Q4_K_M. Options come from file listings on Hugging Face, collected on 2026-10-11: 503 in total. This round lists only files published by the lab itself. Community conversions are not listed yet.
For original weights the size is the total of the weight files in the repository, which can include more than one variant. In GGUF repositories, projector, calibration and auxiliary-head files are not counted as options.
Minimum memory
Memory need belongs to a weight option, not to a model name. A figure is marked ● when the publisher states it and ◇ when it is our estimate. Everything else says Unknown.
Estimates exist only for text language models: the Language models, Coding, Agents and Translation tasks. The estimate is the weight file size, plus the key-value cache for a 8K-token context, plus 0.5 GB for the runtime. The cache per token is 2 × layers × key-value heads × head dimension × 2 bytes, read from the repository's config. When the config does not give those numbers the cache is left out, and the note on the option says so. A longer context needs more memory than the estimate.
Speech, image, video, music and OCR models get no estimate. Their memory use depends on resolution, length and pipeline, and file size is a poor guide. 64 of 154 models have a figure; 90 say Unknown.
Fits, tight, too large
Set your device memory and each option is compared against it. Fits: it needs at most 85% of your memory. Tight: more than 85%, up to 100%. Too large: more than 100%. Unknown: there is no memory figure. This is a memory check only and says nothing about speed.
For a graphics card, enter its video memory. For Apple silicon, enter the unified memory. For CPU only, enter the RAM. MLX files are counted only for Apple silicon. The page detects the device type and shows the graphics chip name as a hint. It never guesses memory. Your settings stay in your browser.
Third-party scores
Some pages quote scores from other evaluators, with the source and the data date each time. They are not this site's measurements. They were run on hosted endpoints, so they do not describe a particular quantized file. Scores from different sources are never put in one column or one chart. 19 models have at least one.
| Source | Data from | Citation |
|---|---|---|
| OpenRouter evals | 2026-10-10 | Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). |
| Artificial Analysis | 2026-10-09 | Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). |
| Design Arena | 2026-10-11 | Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). |
Charts
A chart plots one metric from one source against minimum memory on a log scale. It appears only when at least 8 models have both a score and a memory figure within the chosen memory budget. Otherwise the page shows the table alone. Because the scores are from hosted endpoints, each model is one point, at the memory need of its original weights. The line connects points that no other point beats on both memory and score. It does not predict anything in between.
Tests on this site
This site has not run any model yet, and every model page says so. Current status: Not tested yet 130, Not in the first test round 24.
The first round will use hosted APIs where they exist and rented servers otherwise. Each result will name the endpoint or the hardware. A speed measured there is not a speed on your machine and will not be shown as one.
| Task | Shown side by side | Checked by program |
|---|---|---|
| Language models | An SVG drawing from one fixed prompt | A takeaway order turned into JSON against a schema; a meeting time across time zones; three sentences without the letter e; one sentence found in 20K tokens; a total read off a receipt photo |
| Speech to text | A 30-second café recording, shown against the reference text | Word error rate on our own recordings |
| Text to speech | One paragraph with dates, prices and hard names | Delay to first audio and real-time factor |
| Image generation | A shop sign with fixed wording and objects; an edit that changes one object | The sign text read back by OCR |
| Video and 3D | Five-second clips from fixed prompts | Duration, resolution and frame rate; a closed mesh |
| Music and sound effects | An 8-second jingle at 120 BPM; a 3-second sound effect | Duration, loudness and detected tempo |
| Translation | English to Chinese and Chinese to English | A learned quality score against reference translations |
| OCR and document parsing | A creased handwritten menu with a table | Character error rate and table cells |
Tasks that generate something are run 3 times and every run is shown. Tasks with a checkable answer are run 10 times. The first round does not use a language model as the judge. Code, agent, embedding and safety models are not in the first round; their pages list sources only.