llama.cpp is the C/C++ engine most local-LLM tools are built on. You can use it directly.
1. Install
The official installer (macOS and Linux):
curl -LsSf https://llama.app/install.sh | sh
Installs a single llama binary into ~/.local/bin, which exposes subcommands like llama cli and llama serve.
Package managers work too:
brew install llama.cpp # macOS / Linux
winget install llama.cpp # Windows
Important: package-manager and release-binary builds ship the classic executables llama-cli and llama-server instead of the llama wrapper. Every flag below is identical — only the command name differs. If llama cli is not found, try llama-cli.
Building from source is the third option: clone the repo and follow docs/build.md.
2. Run a model
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
Downloads the GGUF from that Hugging Face repo (cached in ~/.cache/llama.cpp) and starts an interactive chat. -hf defaults to the Q4_K_M quantization; pin another with -hf user/repo:Q8_0.
Already have a .gguf file on disk?
llama cli -m ./models/my-model.gguf -c 8192
Loads that file with an 8192-token context. Drop -c to use whatever the model was built with.
3. Serve an OpenAI-compatible API
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
Starts the HTTP server on 127.0.0.1:8080 with a built-in web UI at http://localhost:8080 and OpenAI-style routes under /v1. Add --port 9000 or --host 0.0.0.0 to change where it listens.
4. Call it
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Name three uses for a raspberry pi."}]}'
Returns a normal OpenAI chat-completion JSON body. Any OpenAI SDK works — set base_url to http://localhost:8080/v1 and pass any non-empty API key.
Useful flags
-ngl N— layers to put on the GPU. Default isauto; set-ngl 0to force CPU-only.-c N— context size. Bigger context costs more RAM/VRAM.-cl— list the models already in your local cache.
Next: choose a GGUF quantization or run the same models through Ollama.