Quantization and weight formats
What the names on a weight file mean, and which current models publish which levels.
Quantization
Models are trained with 16-bit or 32-bit numbers. Quantization rewrites those weights with 8, 4 or even fewer bits. A 4-bit file is roughly a quarter to a third of the 16-bit size, which is what makes many models fit on consumer hardware. The loss in quality is small at 8 bits and grows as bits drop, and it differs by model and by task, so a general rule is a poor guide for any one model.
Sources: Hugging Face: GGUF
GGUF
GGUF stores a model's tensors together with a standard set of metadata, so a runner can load it without extra configuration. Most GGUF files are quantized, and the quantization level is part of the file name, such as Q4_K_M or Q8_0. Large models are often split into several numbered parts. GGUF files run through llama.cpp and the tools built on it, on CPUs and on GPUs.
Sources: Hugging Face: GGUF · llama.cpp
Q4_K_M
In GGUF names, the number is the approximate bit width, K marks llama.cpp's k-quant method, and the last letter (S, M or L) says how many tensors are kept at higher precision. Q4_K_M sits in the middle: much smaller than 8-bit, usually with a modest quality drop. In llama.cpp's own example an 8B model goes from 32.1 GB at 32-bit to 4.9 GB at Q4_K_M. How much quality is lost depends on the model and the task.
Sources: llama.cpp quantize tool · Hugging Face GGUF quantization types
Q8_0
Q8_0 stores each weight in 8 bits, rounded to the nearest value. It is the level to pick when you have the memory and want output as close as possible to the original model while still using a GGUF file.
Sources: Hugging Face GGUF quantization types
MLX
MLX is an array framework from Apple's machine learning research team, designed around the shared memory of Apple silicon. MLX model files are a separate conversion from GGUF, with their own quantization scheme, so an MLX 4-bit file and a GGUF Q4_K_M file of the same model are different files and can behave slightly differently. Results measured on one do not carry over to the other. MLX also has Linux builds, but this site counts MLX files only for Apple silicon.
Sources: MLX
BF16
BF16 and FP16 both use 16 bits per weight and differ in how those bits are split between range and precision. A model's size in gigabytes at 16-bit is about twice its parameter count in billions.
Models with GGUF or MLX files
| Model | Lab | Published levels |
|---|---|---|
| ArmorOCR | Ant Group inclusionAI | GGUF: Q4_K_M, Q8_0 |
| BitCPM-CANN | OpenBMB | GGUF: TQ2_0, BF16 |
| BitNet Embedding | Microsoft | GGUF: BF16 |
| colgrep agent 2B | LightOn | GGUF: Q4_0, Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16, F16 |
| Granite 4.2 | IBM | GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_0, Q4_K_S, Q4_K_M, Q4_1, Q5_0, Q5_K_S, Q5_K_M, Q5_1, Q6_K, Q8_0, BF16 MLX: 4-bit, 8-bit, BF16 |
| Granite Guardian 4.1 | IBM | GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_0, Q4_K_S, Q4_K_M, Q4_1, Q5_0, Q5_K_S, Q5_K_M, Q5_1, Q6_K, Q8_0, BF16 |
| Granite Vision 4.1 | IBM | GGUF: Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16 |
| Hy MT2 | Tencent | GGUF: 1.25-bit, 2-bit, Q4_K_M, Q6_K, Q8_0 |
| Index Echo | Bilibili IndexTeam | GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, F16 |
| Index Homura | Bilibili IndexTeam | GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, F16 |
| Index Translate | Bilibili IndexTeam | GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, F16 |
| Jina Reranker 3.5 | Jina AI | GGUF: IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_XS, IQ3_S, IQ4_XS, IQ4_NL, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, BF16 MLX: unlabeled |
| Ling 3.0 | Ant Group inclusionAI | GGUF: Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16 |
| Mellum2.1 Thinking | JetBrains | GGUF: MXFP4_MOE, Q4_K_M, Q6_K, Q8_0, BF16 |
| MiniCPM-V 4.6 | OpenBMB | GGUF: Q4_0, Q4_K_S, Q4_K_M, Q4_1, Q5_0, Q5_K_S, Q5_K_M, Q5_1, Q6_K, Q8_0, F16 |
| MiniCPM5 | OpenBMB | GGUF: Q4_K_M, Q8_0, F16 MLX: unlabeled |
| Nemotron 3 Diarization | NVIDIA | GGUF: Q8_0 |
| Nemotron 3.5 ASR | NVIDIA | GGUF: Q8_0 |
| PaddleOCR VL 1.6 | Baidu PaddlePaddle | GGUF: unlabeled |
| Qwen 27B Blend | JetBrains | GGUF: IQ3_S, IQ4_XS, Q4_K_M, Q5_K_M MLX: 4-bit |
| SingGuard | Ant Group inclusionAI | GGUF: Q4_K_M, Q8_0, F16 |
| Step 3.7 Flash | StepFun | GGUF: IQ3_XXS, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q8_0, BF16 |
| VibeVoice ASR BitNet | Microsoft | GGUF: unlabeled, Q6_K |
Files published by the lab itself. Community conversions are not listed. Each model page gives the size and memory need of every file.