# llama.cpp Quickstart: Run a GGUF Model

> Install llama.cpp, pull a GGUF straight from Hugging Face with llama cli -hf, chat in the terminal, then serve an OpenAI-compatible API on localhost:8080.

- Canonical: https://guides-ai.pages.dev/guides/llama-cpp-quickstart/
- Plate 09.07 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 06 Sept 2026 · 4 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

llama.cpp is the C/C++ engine most local-LLM tools are built on. You can use it directly.

## 1. Install

The official installer (macOS and Linux):

```bash
curl -LsSf https://llama.app/install.sh | sh
```

Installs a single `llama` binary into `~/.local/bin`, which exposes subcommands like `llama cli` and `llama serve`.

Package managers work too:

```bash
brew install llama.cpp      # macOS / Linux
winget install llama.cpp    # Windows
```

**Important:** package-manager and release-binary builds ship the classic executables `llama-cli` and `llama-server` instead of the `llama` wrapper. Every flag below is identical — only the command name differs. If `llama cli` is not found, try `llama-cli`.

Building from source is the third option: clone the repo and follow `docs/build.md`.

## 2. Run a model

```bash
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
```

Downloads the GGUF from that Hugging Face repo (cached in `~/.cache/llama.cpp`) and starts an interactive chat. `-hf` defaults to the `Q4_K_M` quantization; pin another with `-hf user/repo:Q8_0`.

Already have a `.gguf` file on disk?

```bash
llama cli -m ./models/my-model.gguf -c 8192
```

Loads that file with an 8192-token context. Drop `-c` to use whatever the model was built with.

## 3. Serve an OpenAI-compatible API

```bash
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
```

Starts the HTTP server on `127.0.0.1:8080` with a built-in web UI at `http://localhost:8080` and OpenAI-style routes under `/v1`. Add `--port 9000` or `--host 0.0.0.0` to change where it listens.

## 4. Call it

```bash
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Name three uses for a raspberry pi."}]}'
```

Returns a normal OpenAI chat-completion JSON body. Any OpenAI SDK works — set `base_url` to `http://localhost:8080/v1` and pass any non-empty API key.

## Useful flags

- `-ngl N` — layers to put on the GPU. Default is `auto`; set `-ngl 0` to force CPU-only.
- `-c N` — context size. Bigger context costs more RAM/VRAM.
- `-cl` — list the models already in your local cache.

---

Next: [choose a GGUF quantization](/guides/choose-llm-quantization/) or [run the same models through Ollama](/guides/run-local-llm-ollama-macos-linux/).
