# Choose a GGUF Quantization: Q4 vs Q5 vs Q8

> Q4_K_M is the default pick for local LLMs. Real bits-per-weight for each GGUF level, why Q8_0 is a legacy format, and how to match one to your RAM.

- Canonical: https://guides-ai.pages.dev/guides/choose-llm-quantization/
- Plate 09.04 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 02 Sept 2026 · Updated: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Start at **Q4_K_M**. It is the community default because it keeps most of the model's quality
at roughly a quarter of the 16-bit size, and it is the level most GGUF repos ship first.
Everything below is the reasoning behind that, so you know when to leave it.

## 1. Prefer the K-quants

The suffix matters more than the number. Hugging Face's own quantization table marks `Q4_0`,
`Q5_0`, `Q8_0` and friends as **legacy** round-to-nearest methods "not used widely as of
today". The `_K` types — Q3_K, Q4_K, Q5_K, Q6_K — use super-blocks with per-block scales and
give better quality at the same bit budget. `_M` and `_S` are the medium and small variants
within a level.

Q8_0 is the exception you'll still meet constantly: it is legacy by format but is still what
repos publish as their near-lossless option.

## 2. The levels, with their real cost

```text
Q6_K      6.5625 bits/weight   effectively indistinguishable from the original
Q5_K_M    5.5    bits/weight   safe middle ground
Q4_K_M    4.5    bits/weight   default pick — small, quality drop most people never notice
Q3_K_M    3.4375 bits/weight   noticeably rougher
Q2_K      2.625  bits/weight   last resort; degradation stops being subtle
```

What it does: the exact bits-per-weight from the GGUF spec, which is what actually sets the
file size. Below 4 bits, quality falls off a cliff rather than a slope.

## 3. Turn bits into gigabytes

```text
file size ≈ parameters × bits per weight ÷ 8

8B  × 4.5 (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B  × 8   (Q8_0)   ÷ 8 ≈ 8   GB
```

What it does: predicts the download before you start it, so you can rule a model out in ten
seconds.

## Verify it worked

The test is not the benchmark, it's the speed. Load the model and watch generation: if it
streams at roughly reading pace, it fits. If it crawls out one word per second, it is paging to
disk — drop a level or shrink the context window before blaming the
[quantization](/glossary/#quantization).

The rule that beats all of this: **fit the largest model you can at Q4 before choosing a
smaller model at Q8.** Parameter count buys more than precision does.

Next: [how much RAM or VRAM a local LLM needs](/guides/local-llm-ram-vram-requirements/), or
run one with [LM Studio](/guides/run-llm-lm-studio/), which shows the quant level and whether
it fits before you download.

Source: [GGUF quantization types on the Hugging Face Hub](https://huggingface.co/docs/hub/gguf).
