How Much RAM or VRAM a Local LLM Needs (1B to 70B)
Real GGUF file sizes at Q4/Q8 for 1B–70B models, the rule of thumb behind them, and what fits in 8, 16, 24 or 64 GB. Pick a model that actually runs.
On this page5 sections
The model file has to fit in memory — GPU memory (VRAM) if you want speed, system RAM if you’re on CPU or an Apple-silicon Mac, where both are the same pool. Everything below is measured from real files, not marketing.
1. The rule of thumb
A model’s size on disk is roughly parameters × bits per weight ÷ 8, plus a few percent of overhead:
8B × 4.5 bits (Q4_K_M) ÷ 8 ≈ 4.5 GB
8B × 8 bits (Q8_0) ÷ 8 ≈ 8 GB
70B × 4.5 bits (Q4_K_M) ÷ 8 ≈ 40 GB
Q4_K_M is the usual sweet spot: about half the size of Q8 with a small quality loss. Q8_0 is close to lossless. Below Q4 (Q3, Q2) quality drops fast — use it only when nothing else fits.
2. Real sizes, measured
File sizes from the published GGUF repos (bartowski) and Ollama’s library tags:
model Q4_K_M Q8_0 source
gemma3:1b (Ollama) 0.8 GB — ollama.com/library/gemma3
gemma3:4b (Ollama) 3.3 GB — ollama.com/library/gemma3
Llama 3.1 8B 4.6 GB 8.0 GB bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
Qwen3 32B 18.4 GB 32.4 GB bartowski/Qwen_Qwen3-32B-GGUF
Llama 3.3 70B 39.6 GB ~75 GB bartowski/Llama-3.3-70B-Instruct-GGUF
That last column is what ollama pull or hf download actually fetches; the Q8 70B figure is the rule of thumb because that file ships split into parts.
3. Add headroom for context
The file isn’t the whole story. The KV cache (the model’s working memory for your conversation) grows with context length, and the runtime keeps a little for itself. Budget:
total ≈ file size + 1–2 GB (runtime) + KV cache
KV cache: ~0.5 GB per 4k tokens for an 8B model; several GB at 32k+ tokens
Long contexts (num_ctx 32k and up) are where a model that “fits” starts swapping. If generation suddenly crawls, shrink the context before you shrink the model.
4. What fits where
Aim for the file size plus about 20 % headroom:
8 GB VRAM / RAM 1B–4B at Q4–Q8, 7–8B at Q4 with a short context
16 GB 8B at Q8, or 12–14B at Q4
24 GB (RTX 4090) 32B at Q4 fits tightly (18.4 GB) — keep context modest
32 GB Mac 32B at Q4 comfortably, 8B at Q8 with long context
64 GB Mac 70B at Q4 (39.6 GB) with room for context
On a Mac, unified memory is shared with everything else you have open — close the browser tabs before loading a 70B.
5. Check before you pull
The tag page tells you the exact size:
ollama show llama3.1:8b --modelfile # after a pull: shows quantization and params
Or read the file list on the Hugging Face repo before hf download — each GGUF shows its size next to the quantization name. When a model runs but streams at one word a second, it’s paging: pick the next quantization down or a smaller model.
Next: choose a GGUF quantization · run a local LLM with Ollama on macOS and Linux.