Local LLMs

How Much RAM or VRAM a Local LLM Needs (1B to 70B)

3 min read

The model file has to fit in memory — GPU memory (VRAM) if you want speed, system RAM if you’re on CPU or an Apple-silicon Mac, where both are the same pool. Everything below is measured from real files, not marketing.

1. The rule of thumb

A model’s size on disk is roughly parameters × bits per weight ÷ 8, plus a few percent of overhead:

8B  × 4.5 bits (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B  × 8   bits (Q8_0)   ÷ 8 ≈ 8   GB
70B × 4.5 bits (Q4_K_M) ÷ 8 ≈ 40  GB

Q4_K_M is the usual sweet spot: about half the size of Q8 with a small quality loss. Q8_0 is close to lossless. Below Q4 (Q3, Q2) quality drops fast — use it only when nothing else fits.

2. Real sizes, measured

File sizes from the published GGUF repos (bartowski) and Ollama’s library tags:

model                     Q4_K_M    Q8_0     source
gemma3:1b (Ollama)         0.8 GB   —        ollama.com/library/gemma3
gemma3:4b (Ollama)         3.3 GB   —        ollama.com/library/gemma3
Llama 3.1 8B               4.6 GB   8.0 GB   bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
Qwen3 32B                 18.4 GB  32.4 GB   bartowski/Qwen_Qwen3-32B-GGUF
Llama 3.3 70B             39.6 GB   ~75 GB   bartowski/Llama-3.3-70B-Instruct-GGUF

That last column is what ollama pull or hf download actually fetches; the Q8 70B figure is the rule of thumb because that file ships split into parts.

3. Add headroom for context

The file isn’t the whole story. The KV cache (the model’s working memory for your conversation) grows with context length, and the runtime keeps a little for itself. Budget:

total ≈ file size + 1–2 GB (runtime) + KV cache
KV cache: ~0.5 GB per 4k tokens for an 8B model; several GB at 32k+ tokens

Long contexts (num_ctx 32k and up) are where a model that “fits” starts swapping. If generation suddenly crawls, shrink the context before you shrink the model.

4. What fits where

Aim for the file size plus about 20 % headroom:

8 GB VRAM / RAM     1B–4B at Q4–Q8, 7–8B at Q4 with a short context
16 GB               8B at Q8, or 12–14B at Q4
24 GB (RTX 4090)    32B at Q4 fits tightly (18.4 GB) — keep context modest
32 GB Mac           32B at Q4 comfortably, 8B at Q8 with long context
64 GB Mac           70B at Q4 (39.6 GB) with room for context

On a Mac, unified memory is shared with everything else you have open — close the browser tabs before loading a 70B.

5. Check before you pull

The tag page tells you the exact size:

ollama show llama3.1:8b --modelfile   # after a pull: shows quantization and params

Or read the file list on the Hugging Face repo before hf download — each GGUF shows its size next to the quantization name. When a model runs but streams at one word a second, it’s paging: pick the next quantization down or a smaller model.


Next: choose a GGUF quantization · run a local LLM with Ollama on macOS and Linux.

Open the full interactive version (with copy buttons) ↗

← All guides