§09.07

llama.cpp Quickstart: Run a GGUF Model

Install llama.cpp, pull a GGUF straight from Hugging Face with llama cli -hf, chat in the terminal, then serve an OpenAI-compatible API on localhost:8080.

published 06 Sept 2026 checked against docs 06 Sept 2026 4 min in Local LLMs Markdown

Step 2 of 4 · Serve models yourself

On this page5 sections
  1. 1. Install
  2. 2. Run a model
  3. 3. Serve an OpenAI-compatible API
  4. 4. Call it
  5. Useful flags

llama.cpp is the C/C++ engine most local-LLM tools are built on. You can use it directly.

1. Install

The official installer (macOS and Linux):

curl -LsSf https://llama.app/install.sh | sh

Installs a single llama binary into ~/.local/bin, which exposes subcommands like llama cli and llama serve.

Package managers work too:

brew install llama.cpp      # macOS / Linux
winget install llama.cpp    # Windows

Important: package-manager and release-binary builds ship the classic executables llama-cli and llama-server instead of the llama wrapper. Every flag below is identical — only the command name differs. If llama cli is not found, try llama-cli.

Building from source is the third option: clone the repo and follow docs/build.md.

2. Run a model

llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

Downloads the GGUF from that Hugging Face repo (cached in ~/.cache/llama.cpp) and starts an interactive chat. -hf defaults to the Q4_K_M quantization; pin another with -hf user/repo:Q8_0.

Already have a .gguf file on disk?

llama cli -m ./models/my-model.gguf -c 8192

Loads that file with an 8192-token context. Drop -c to use whatever the model was built with.

3. Serve an OpenAI-compatible API

llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

Starts the HTTP server on 127.0.0.1:8080 with a built-in web UI at http://localhost:8080 and OpenAI-style routes under /v1. Add --port 9000 or --host 0.0.0.0 to change where it listens.

4. Call it

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Name three uses for a raspberry pi."}]}'

Returns a normal OpenAI chat-completion JSON body. Any OpenAI SDK works — set base_url to http://localhost:8080/v1 and pass any non-empty API key.

Useful flags

  • -ngl N — layers to put on the GPU. Default is auto; set -ngl 0 to force CPU-only.
  • -c N — context size. Bigger context costs more RAM/VRAM.
  • -cl — list the models already in your local cache.

Next: choose a GGUF quantization or run the same models through Ollama.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list