Local LLMs

Choose a GGUF Quantization: Q4 vs Q5 vs Q8

3 min read

Start at Q4_K_M. It is the community default because it keeps most of the model’s quality at roughly a quarter of the 16-bit size, and it is the level most GGUF repos ship first. Everything below is the reasoning behind that, so you know when to leave it.

1. Prefer the K-quants

The suffix matters more than the number. Hugging Face’s own quantization table marks Q4_0, Q5_0, Q8_0 and friends as legacy round-to-nearest methods “not used widely as of today”. The _K types — Q3_K, Q4_K, Q5_K, Q6_K — use super-blocks with per-block scales and give better quality at the same bit budget. _M and _S are the medium and small variants within a level.

Q8_0 is the exception you’ll still meet constantly: it is legacy by format but is still what repos publish as their near-lossless option.

2. The levels, with their real cost

Q6_K      6.5625 bits/weight   effectively indistinguishable from the original
Q5_K_M    5.5    bits/weight   safe middle ground
Q4_K_M    4.5    bits/weight   default pick — small, quality drop most people never notice
Q3_K_M    3.4375 bits/weight   noticeably rougher
Q2_K      2.625  bits/weight   last resort; degradation stops being subtle

What it does: the exact bits-per-weight from the GGUF spec, which is what actually sets the file size. Below 4 bits, quality falls off a cliff rather than a slope.

3. Turn bits into gigabytes

file size ≈ parameters × bits per weight ÷ 8

8B  × 4.5 (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B  × 8   (Q8_0)   ÷ 8 ≈ 8   GB

What it does: predicts the download before you start it, so you can rule a model out in ten seconds.

Verify it worked

The test is not the benchmark, it’s the speed. Load the model and watch generation: if it streams at roughly reading pace, it fits. If it crawls out one word per second, it is paging to disk — drop a level or shrink the context window before blaming the quantization.

The rule that beats all of this: fit the largest model you can at Q4 before choosing a smaller model at Q8. Parameter count buys more than precision does.

Next: how much RAM or VRAM a local LLM needs, or run one with LM Studio, which shows the quant level and whether it fits before you download.

Source: GGUF quantization types on the Hugging Face Hub.

Open the full interactive version (with copy buttons) ↗

← All guides