Choose a GGUF Quantization: Q4 vs Q5 vs Q8
Q4_K_M is the default pick for local LLMs. Real bits-per-weight for each GGUF level, why Q8_0 is a legacy format, and how to match one to your RAM.
Step 3 of 5 · Run models locally
On this page4 sections
Start at Q4_K_M. It is the community default because it keeps most of the model’s quality at roughly a quarter of the 16-bit size, and it is the level most GGUF repos ship first. Everything below is the reasoning behind that, so you know when to leave it.
1. Prefer the K-quants
The suffix matters more than the number. Hugging Face’s own quantization table marks Q4_0,
Q5_0, Q8_0 and friends as legacy round-to-nearest methods “not used widely as of
today”. The _K types — Q3_K, Q4_K, Q5_K, Q6_K — use super-blocks with per-block scales and
give better quality at the same bit budget. _M and _S are the medium and small variants
within a level.
Q8_0 is the exception you’ll still meet constantly: it is legacy by format but is still what repos publish as their near-lossless option.
2. The levels, with their real cost
Q6_K 6.5625 bits/weight effectively indistinguishable from the original
Q5_K_M 5.5 bits/weight safe middle ground
Q4_K_M 4.5 bits/weight default pick — small, quality drop most people never notice
Q3_K_M 3.4375 bits/weight noticeably rougher
Q2_K 2.625 bits/weight last resort; degradation stops being subtle
What it does: the exact bits-per-weight from the GGUF spec, which is what actually sets the file size. Below 4 bits, quality falls off a cliff rather than a slope.
3. Turn bits into gigabytes
file size ≈ parameters × bits per weight ÷ 8
8B × 4.5 (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B × 8 (Q8_0) ÷ 8 ≈ 8 GB
What it does: predicts the download before you start it, so you can rule a model out in ten seconds.
Verify it worked
The test is not the benchmark, it’s the speed. Load the model and watch generation: if it streams at roughly reading pace, it fits. If it crawls out one word per second, it is paging to disk — drop a level or shrink the context window before blaming the quantization.
The rule that beats all of this: fit the largest model you can at Q4 before choosing a smaller model at Q8. Parameter count buys more than precision does.
Next: how much RAM or VRAM a local LLM needs, or run one with LM Studio, which shows the quant level and whether it fits before you download.