# Download Models from Hugging Face with the hf CLI

> Install the hf CLI, log in, download one GGUF instead of a whole repo with --include, put files where you want with --local-dir, and clean the cache.

- Canonical: https://guides-ai.pages.dev/guides/hugging-face-cli-download-models/
- Plate 09.06 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Every local-model tool eventually needs a file from the Hub — a GGUF for llama.cpp, a diffusion checkpoint, an adapter. The `hf` CLI (the successor of `huggingface-cli`, from the `huggingface_hub` package) downloads exactly the files you want, resumes interrupted transfers, and manages the cache.

## 1. Install

```bash
curl -LsSf https://hf.co/cli/install.sh | bash
```

```powershell
powershell -ExecutionPolicy ByPass -c "irm https://hf.co/cli/install.ps1 | iex"
```

The standalone installer puts `hf` on your PATH (and adds an `hf-cli` skill for coding agents; add `--exclude-skill` to skip that). Already have Python? `uvx hf download …` runs it without installing. Check with `hf --help`.

## 2. Log in (needed for gated repos)

```bash
hf auth login
```

Paste a token from your Hugging Face settings. Gated models (Llama, FLUX, …) also need you to accept the licence on the model page once; without both you get a 401 on download. In CI use `hf auth login --token $HF_TOKEN`.

## 3. Download one file, not 60 GB

Whole repo, into the cache:

```bash
hf download HuggingFaceH4/zephyr-7b-beta
```

One file — the usual case for a GGUF:

```bash
hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf
```

A filtered set with glob patterns:

```bash
hf download stabilityai/stable-diffusion-xl-base-1.0 --include "*.safetensors" --exclude "*.fp16.*"
```

`--include`/`--exclude` take shell-style patterns; datasets need `--repo-type dataset`; a specific tag or commit is `--revision v1.1`. Add `--dry-run` to see what would be fetched and how big it is.

## 4. Cache or a folder you choose

By default files land in the cache under `HF_HOME` (`~/.cache/huggingface` unless you set the variable) and the command prints the path — pass that path to llama.cpp or Ollama's `FROM`. To get a plain folder instead:

```bash
hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf --local-dir ./models
```

`--local-dir` writes real files where you point it (good for a project folder); the cache is better when several tools share the same model. `--quiet` prints only the final path, handy in scripts.

## 5. Keep the disk in check

```bash
hf cache ls
hf cache ls --filter "size>30g" --revisions
hf cache rm model/gpt2
hf cache prune
```

`ls` shows what's stored and how big; `rm` removes a repo or a revision; `prune` deletes `.incomplete` leftovers from interrupted downloads. To delete everything you haven't touched in a year: `hf cache rm $(hf cache ls --filter "accessed>1y" -q) -y`.

---

Next: [llama.cpp quickstart](/guides/llama-cpp-quickstart/) · [how much RAM or VRAM a model needs](/guides/local-llm-ram-vram-requirements/) · [choose a GGUF quantization](/guides/choose-llm-quantization/).
