Ollama runs open models on your own machine. Same install command on macOS and Linux, no API key.
1. Install
curl -fsSL https://ollama.com/install.sh | sh
On macOS this installs Ollama.app into /Applications and links the ollama command into /usr/local/bin. On Linux it installs the binary and writes a systemd unit at /etc/systemd/system/ollama.service.
Check it worked:
ollama -v
Prints the installed version — if the command is not found, open a new terminal so your PATH refreshes.
2. Pull a small model
ollama pull gemma3:4b
Downloads a ~3.3 GB 4B model. Start here rather than with the default gemma4, which is a ~9.6 GB download.
3. Chat
ollama run gemma3:4b
Loads the model and drops you into a prompt. Type /bye to leave the chat and unload it.
4. Keep it running as a service
On Linux, the installer already enabled the service. To check, control, and watch it:
systemctl status ollama
sudo systemctl enable --now ollama
journalctl -u ollama -f
Shows whether the daemon is up, makes it start at boot, and tails its log. The API is then always available on http://localhost:11434.
On macOS, the app starts at login and does the same job. If you would rather run it manually in a terminal:
ollama serve
Runs the daemon in the foreground; leave that window open while you use the API.
Choosing a size for your RAM
You need roughly the download size in free RAM (or VRAM), plus 1–2 GB of headroom for context:
| Command | Download | Comfortable on |
|---|---|---|
ollama run gemma3:1b | ~815 MB | 4 GB RAM |
ollama run qwen3.5:2b | ~2.7 GB | 8 GB RAM |
ollama run gemma3:4b | ~3.3 GB | 8 GB RAM |
ollama run gemma4 | ~9.6 GB | 16 GB RAM |
If a model swaps to disk it still answers, just very slowly — that is the signal to drop a size.
Housekeeping
ollama list
ollama ps
ollama rm gemma3:4b
Lists what you downloaded, shows what is loaded in memory right now, and deletes a model you no longer want.
Next: call your local model from Python and curl or pick a GGUF quantization.