Quantization and weight formats

What the names on a weight file mean, and which current models publish which levels.

Quantization

Models are trained with 16-bit or 32-bit numbers. Quantization rewrites those weights with 8, 4 or even fewer bits. A 4-bit file is roughly a quarter to a third of the 16-bit size, which is what makes many models fit on consumer hardware. The loss in quality is small at 8 bits and grows as bits drop, and it differs by model and by task, so a general rule is a poor guide for any one model.

Sources: Hugging Face: GGUF

GGUF

GGUF stores a model's tensors together with a standard set of metadata, so a runner can load it without extra configuration. Most GGUF files are quantized, and the quantization level is part of the file name, such as Q4_K_M or Q8_0. Large models are often split into several numbered parts. GGUF files run through llama.cpp and the tools built on it, on CPUs and on GPUs.

Sources: Hugging Face: GGUF · llama.cpp

Q4_K_M

In GGUF names, the number is the approximate bit width, K marks llama.cpp's k-quant method, and the last letter (S, M or L) says how many tensors are kept at higher precision. Q4_K_M sits in the middle: much smaller than 8-bit, usually with a modest quality drop. In llama.cpp's own example an 8B model goes from 32.1 GB at 32-bit to 4.9 GB at Q4_K_M. How much quality is lost depends on the model and the task.

Sources: llama.cpp quantize tool · Hugging Face GGUF quantization types

Q8_0

Q8_0 stores each weight in 8 bits, rounded to the nearest value. It is the level to pick when you have the memory and want output as close as possible to the original model while still using a GGUF file.

Sources: Hugging Face GGUF quantization types

MLX

MLX is an array framework from Apple's machine learning research team, designed around the shared memory of Apple silicon. MLX model files are a separate conversion from GGUF, with their own quantization scheme, so an MLX 4-bit file and a GGUF Q4_K_M file of the same model are different files and can behave slightly differently. Results measured on one do not carry over to the other. MLX also has Linux builds, but this site counts MLX files only for Apple silicon.

Sources: MLX

BF16

BF16 and FP16 both use 16 bits per weight and differ in how those bits are split between range and precision. A model's size in gigabytes at 16-bit is about twice its parameter count in billions.

Models with GGUF or MLX files

23 models · collected 2026-10-11
ModelLabPublished levels
ArmorOCRAnt Group inclusionAI
GGUF: Q4_K_M, Q8_0
BitCPM-CANNOpenBMB
GGUF: TQ2_0, BF16
BitNet EmbeddingMicrosoft
GGUF: BF16
colgrep agent 2BLightOn
GGUF: Q4_0, Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16, F16
Granite 4.2IBM
GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_0, Q4_K_S, Q4_K_M, Q4_1, Q5_0, Q5_K_S, Q5_K_M, Q5_1, Q6_K, Q8_0, BF16
MLX: 4-bit, 8-bit, BF16
Granite Guardian 4.1IBM
GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_0, Q4_K_S, Q4_K_M, Q4_1, Q5_0, Q5_K_S, Q5_K_M, Q5_1, Q6_K, Q8_0, BF16
Granite Vision 4.1IBM
GGUF: Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16
Hy MT2Tencent
GGUF: 1.25-bit, 2-bit, Q4_K_M, Q6_K, Q8_0
Index EchoBilibili IndexTeam
GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, F16
Index HomuraBilibili IndexTeam
GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, F16
Index TranslateBilibili IndexTeam
GGUF: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, F16
Jina Reranker 3.5Jina AI
GGUF: IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_XS, IQ3_S, IQ4_XS, IQ4_NL, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, BF16
MLX: unlabeled
Ling 3.0Ant Group inclusionAI
GGUF: Q4_K_M, Q5_K_M, Q6_K, Q8_0, BF16
Mellum2.1 ThinkingJetBrains
GGUF: MXFP4_MOE, Q4_K_M, Q6_K, Q8_0, BF16
MiniCPM-V 4.6OpenBMB
GGUF: Q4_0, Q4_K_S, Q4_K_M, Q4_1, Q5_0, Q5_K_S, Q5_K_M, Q5_1, Q6_K, Q8_0, F16
MiniCPM5OpenBMB
GGUF: Q4_K_M, Q8_0, F16
MLX: unlabeled
Nemotron 3 DiarizationNVIDIA
GGUF: Q8_0
Nemotron 3.5 ASRNVIDIA
GGUF: Q8_0
PaddleOCR VL 1.6Baidu PaddlePaddle
GGUF: unlabeled
Qwen 27B BlendJetBrains
GGUF: IQ3_S, IQ4_XS, Q4_K_M, Q5_K_M
MLX: 4-bit
SingGuardAnt Group inclusionAI
GGUF: Q4_K_M, Q8_0, F16
Step 3.7 FlashStepFun
GGUF: IQ3_XXS, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S, Q8_0, BF16
VibeVoice ASR BitNetMicrosoft
GGUF: unlabeled, Q6_K

Files published by the lab itself. Community conversions are not listed. Each model page gives the size and memory need of every file.