§09.04

Choose a GGUF Quantization: Q4 vs Q5 vs Q8

Q4_K_M is the default pick for local LLMs. Real bits-per-weight for each GGUF level, why Q8_0 is a legacy format, and how to match one to your RAM.

published 02 Sept 2026 updated 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Local LLMs Markdown

Step 3 of 5 · Run models locally

On this page4 sections
  1. 1. Prefer the K-quants
  2. 2. The levels, with their real cost
  3. 3. Turn bits into gigabytes
  4. Verify it worked

Start at Q4_K_M. It is the community default because it keeps most of the model’s quality at roughly a quarter of the 16-bit size, and it is the level most GGUF repos ship first. Everything below is the reasoning behind that, so you know when to leave it.

1. Prefer the K-quants

The suffix matters more than the number. Hugging Face’s own quantization table marks Q4_0, Q5_0, Q8_0 and friends as legacy round-to-nearest methods “not used widely as of today”. The _K types — Q3_K, Q4_K, Q5_K, Q6_K — use super-blocks with per-block scales and give better quality at the same bit budget. _M and _S are the medium and small variants within a level.

Q8_0 is the exception you’ll still meet constantly: it is legacy by format but is still what repos publish as their near-lossless option.

2. The levels, with their real cost

Q6_K      6.5625 bits/weight   effectively indistinguishable from the original
Q5_K_M    5.5    bits/weight   safe middle ground
Q4_K_M    4.5    bits/weight   default pick — small, quality drop most people never notice
Q3_K_M    3.4375 bits/weight   noticeably rougher
Q2_K      2.625  bits/weight   last resort; degradation stops being subtle

What it does: the exact bits-per-weight from the GGUF spec, which is what actually sets the file size. Below 4 bits, quality falls off a cliff rather than a slope.

3. Turn bits into gigabytes

file size ≈ parameters × bits per weight ÷ 8

8B  × 4.5 (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B  × 8   (Q8_0)   ÷ 8 ≈ 8   GB

What it does: predicts the download before you start it, so you can rule a model out in ten seconds.

Verify it worked

The test is not the benchmark, it’s the speed. Load the model and watch generation: if it streams at roughly reading pace, it fits. If it crawls out one word per second, it is paging to disk — drop a level or shrink the context window before blaming the quantization.

The rule that beats all of this: fit the largest model you can at Q4 before choosing a smaller model at Q8. Parameter count buys more than precision does.

Next: how much RAM or VRAM a local LLM needs, or run one with LM Studio, which shows the quant level and whether it fits before you download.

Source: GGUF quantization types on the Hugging Face Hub.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list