Local LLMs

llama.cpp Quickstart: Run a GGUF Model

4 min read

llama.cpp is the C/C++ engine most local-LLM tools are built on. You can use it directly.

1. Install

The official installer (macOS and Linux):

curl -LsSf https://llama.app/install.sh | sh

Installs a single llama binary into ~/.local/bin, which exposes subcommands like llama cli and llama serve.

Package managers work too:

brew install llama.cpp      # macOS / Linux
winget install llama.cpp    # Windows

Important: package-manager and release-binary builds ship the classic executables llama-cli and llama-server instead of the llama wrapper. Every flag below is identical — only the command name differs. If llama cli is not found, try llama-cli.

Building from source is the third option: clone the repo and follow docs/build.md.

2. Run a model

llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

Downloads the GGUF from that Hugging Face repo (cached in ~/.cache/llama.cpp) and starts an interactive chat. -hf defaults to the Q4_K_M quantization; pin another with -hf user/repo:Q8_0.

Already have a .gguf file on disk?

llama cli -m ./models/my-model.gguf -c 8192

Loads that file with an 8192-token context. Drop -c to use whatever the model was built with.

3. Serve an OpenAI-compatible API

llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

Starts the HTTP server on 127.0.0.1:8080 with a built-in web UI at http://localhost:8080 and OpenAI-style routes under /v1. Add --port 9000 or --host 0.0.0.0 to change where it listens.

4. Call it

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Name three uses for a raspberry pi."}]}'

Returns a normal OpenAI chat-completion JSON body. Any OpenAI SDK works — set base_url to http://localhost:8080/v1 and pass any non-empty API key.

Useful flags


Next: choose a GGUF quantization or run the same models through Ollama.

Open the full interactive version (with copy buttons) ↗

← All guides