Local LLMs

Download Models from Hugging Face with the hf CLI

3 min read

Every local-model tool eventually needs a file from the Hub — a GGUF for llama.cpp, a diffusion checkpoint, an adapter. The hf CLI (the successor of huggingface-cli, from the huggingface_hub package) downloads exactly the files you want, resumes interrupted transfers, and manages the cache.

1. Install

curl -LsSf https://hf.co/cli/install.sh | bash
powershell -ExecutionPolicy ByPass -c "irm https://hf.co/cli/install.ps1 | iex"

The standalone installer puts hf on your PATH (and adds an hf-cli skill for coding agents; add --exclude-skill to skip that). Already have Python? uvx hf download … runs it without installing. Check with hf --help.

2. Log in (needed for gated repos)

hf auth login

Paste a token from your Hugging Face settings. Gated models (Llama, FLUX, …) also need you to accept the licence on the model page once; without both you get a 401 on download. In CI use hf auth login --token $HF_TOKEN.

3. Download one file, not 60 GB

Whole repo, into the cache:

hf download HuggingFaceH4/zephyr-7b-beta

One file — the usual case for a GGUF:

hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf

A filtered set with glob patterns:

hf download stabilityai/stable-diffusion-xl-base-1.0 --include "*.safetensors" --exclude "*.fp16.*"

--include/--exclude take shell-style patterns; datasets need --repo-type dataset; a specific tag or commit is --revision v1.1. Add --dry-run to see what would be fetched and how big it is.

4. Cache or a folder you choose

By default files land in the cache under HF_HOME (~/.cache/huggingface unless you set the variable) and the command prints the path — pass that path to llama.cpp or Ollama’s FROM. To get a plain folder instead:

hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf --local-dir ./models

--local-dir writes real files where you point it (good for a project folder); the cache is better when several tools share the same model. --quiet prints only the final path, handy in scripts.

5. Keep the disk in check

hf cache ls
hf cache ls --filter "size>30g" --revisions
hf cache rm model/gpt2
hf cache prune

ls shows what’s stored and how big; rm removes a repo or a revision; prune deletes .incomplete leftovers from interrupted downloads. To delete everything you haven’t touched in a year: hf cache rm $(hf cache ls --filter "accessed>1y" -q) -y.


Next: llama.cpp quickstart · how much RAM or VRAM a model needs · choose a GGUF quantization.

Open the full interactive version (with copy buttons) ↗

← All guides