§09.06

Download Models from Hugging Face with the hf CLI

Install the hf CLI, log in, download one GGUF instead of a whole repo with --include, put files where you want with --local-dir, and clean the cache.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Local LLMs Markdown

On this page5 sections
  1. 1. Install
  2. 2. Log in (needed for gated repos)
  3. 3. Download one file, not 60 GB
  4. 4. Cache or a folder you choose
  5. 5. Keep the disk in check

Every local-model tool eventually needs a file from the Hub — a GGUF for llama.cpp, a diffusion checkpoint, an adapter. The hf CLI (the successor of huggingface-cli, from the huggingface_hub package) downloads exactly the files you want, resumes interrupted transfers, and manages the cache.

1. Install

curl -LsSf https://hf.co/cli/install.sh | bash
powershell -ExecutionPolicy ByPass -c "irm https://hf.co/cli/install.ps1 | iex"

The standalone installer puts hf on your PATH (and adds an hf-cli skill for coding agents; add --exclude-skill to skip that). Already have Python? uvx hf download … runs it without installing. Check with hf --help.

2. Log in (needed for gated repos)

hf auth login

Paste a token from your Hugging Face settings. Gated models (Llama, FLUX, …) also need you to accept the licence on the model page once; without both you get a 401 on download. In CI use hf auth login --token $HF_TOKEN.

3. Download one file, not 60 GB

Whole repo, into the cache:

hf download HuggingFaceH4/zephyr-7b-beta

One file — the usual case for a GGUF:

hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf

A filtered set with glob patterns:

hf download stabilityai/stable-diffusion-xl-base-1.0 --include "*.safetensors" --exclude "*.fp16.*"

--include/--exclude take shell-style patterns; datasets need --repo-type dataset; a specific tag or commit is --revision v1.1. Add --dry-run to see what would be fetched and how big it is.

4. Cache or a folder you choose

By default files land in the cache under HF_HOME (~/.cache/huggingface unless you set the variable) and the command prints the path — pass that path to llama.cpp or Ollama’s FROM. To get a plain folder instead:

hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf --local-dir ./models

--local-dir writes real files where you point it (good for a project folder); the cache is better when several tools share the same model. --quiet prints only the final path, handy in scripts.

5. Keep the disk in check

hf cache ls
hf cache ls --filter "size>30g" --revisions
hf cache rm model/gpt2
hf cache prune

ls shows what’s stored and how big; rm removes a repo or a revision; prune deletes .incomplete leftovers from interrupted downloads. To delete everything you haven’t touched in a year: hf cache rm $(hf cache ls --filter "accessed>1y" -q) -y.


Next: llama.cpp quickstart · how much RAM or VRAM a model needs · choose a GGUF quantization.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list