# How Much RAM or VRAM a Local LLM Needs (1B to 70B)

> Real GGUF file sizes at Q4/Q8 for 1B–70B models, the rule of thumb behind them, and what fits in 8, 16, 24 or 64 GB. Pick a model that actually runs.

- Canonical: https://guides-ai.pages.dev/guides/local-llm-ram-vram-requirements/
- Plate 09.08 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

The model file has to fit in memory — GPU memory (VRAM) if you want speed, system RAM if you're on CPU or an Apple-silicon Mac, where both are the same pool. Everything below is measured from real files, not marketing.

## 1. The rule of thumb

A model's size on disk is roughly **parameters × bits per weight ÷ 8**, plus a few percent of overhead:

```text
8B  × 4.5 bits (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B  × 8   bits (Q8_0)   ÷ 8 ≈ 8   GB
70B × 4.5 bits (Q4_K_M) ÷ 8 ≈ 40  GB
```

Q4_K_M is the usual sweet spot: about half the size of Q8 with a small quality loss. Q8_0 is close to lossless. Below Q4 (Q3, Q2) quality drops fast — use it only when nothing else fits.

## 2. Real sizes, measured

File sizes from the published GGUF repos (bartowski) and Ollama's library tags:

```text
model                     Q4_K_M    Q8_0     source
gemma3:1b (Ollama)         0.8 GB   —        ollama.com/library/gemma3
gemma3:4b (Ollama)         3.3 GB   —        ollama.com/library/gemma3
Llama 3.1 8B               4.6 GB   8.0 GB   bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
Qwen3 32B                 18.4 GB  32.4 GB   bartowski/Qwen_Qwen3-32B-GGUF
Llama 3.3 70B             39.6 GB   ~75 GB   bartowski/Llama-3.3-70B-Instruct-GGUF
```

That last column is what `ollama pull` or `hf download` actually fetches; the Q8 70B figure is the rule of thumb because that file ships split into parts.

## 3. Add headroom for context

The file isn't the whole story. The KV cache (the model's working memory for your conversation) grows with context length, and the runtime keeps a little for itself. Budget:

```text
total ≈ file size + 1–2 GB (runtime) + KV cache
KV cache: ~0.5 GB per 4k tokens for an 8B model; several GB at 32k+ tokens
```

Long contexts (`num_ctx` 32k and up) are where a model that "fits" starts swapping. If generation suddenly crawls, shrink the context before you shrink the model.

## 4. What fits where

Aim for the file size plus about 20 % headroom:

```text
8 GB VRAM / RAM     1B–4B at Q4–Q8, 7–8B at Q4 with a short context
16 GB               8B at Q8, or 12–14B at Q4
24 GB (RTX 4090)    32B at Q4 fits tightly (18.4 GB) — keep context modest
32 GB Mac           32B at Q4 comfortably, 8B at Q8 with long context
64 GB Mac           70B at Q4 (39.6 GB) with room for context
```

On a Mac, unified memory is shared with everything else you have open — close the browser tabs before loading a 70B.

## 5. Check before you pull

The tag page tells you the exact size:

```bash
ollama show llama3.1:8b --modelfile   # after a pull: shows quantization and params
```

Or read the file list on the Hugging Face repo before `hf download` — each GGUF shows its size next to the quantization name. When a model runs but streams at one word a second, it's paging: pick the next quantization down or a smaller model.

---

Next: [choose a GGUF quantization](/guides/choose-llm-quantization/) · [run a local LLM with Ollama on macOS and Linux](/guides/run-local-llm-ollama-macos-linux/).
