§09.11

How to Run a Local LLM with Ollama on macOS/Linux

Install Ollama with one command, pull gemma3:4b, chat offline in the terminal, and keep the daemon running as a background service on port 11434.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Local LLMs Markdown

Step 1 of 4 · Serve models yourself

On this page6 sections
  1. 1. Install
  2. 2. Pull a small model
  3. 3. Chat
  4. 4. Keep it running as a service
  5. Choosing a size for your RAM
  6. Housekeeping

Ollama runs open models on your own machine. Same install command on macOS and Linux, no API key.

1. Install

curl -fsSL https://ollama.com/install.sh | sh

On macOS this installs Ollama.app into /Applications and links the ollama command into /usr/local/bin. On Linux it installs the binary and writes a systemd unit at /etc/systemd/system/ollama.service.

Check it worked:

ollama -v

Prints the installed version — if the command is not found, open a new terminal so your PATH refreshes.

2. Pull a small model

ollama pull gemma3:4b

Downloads a ~3.3 GB 4B model. Start here rather than with the default gemma4, which is a ~9.6 GB download.

3. Chat

ollama run gemma3:4b

Loads the model and drops you into a prompt. Type /bye to leave the chat and unload it.

4. Keep it running as a service

On Linux, the installer already enabled the service. To check, control, and watch it:

systemctl status ollama
sudo systemctl enable --now ollama
journalctl -u ollama -f

Shows whether the daemon is up, makes it start at boot, and tails its log. The API is then always available on http://localhost:11434.

On macOS, the app starts at login and does the same job. If you would rather run it manually in a terminal:

ollama serve

Runs the daemon in the foreground; leave that window open while you use the API.

Choosing a size for your RAM

You need roughly the download size in free RAM (or VRAM), plus 1–2 GB of headroom for context:

CommandDownloadComfortable on
ollama run gemma3:1b~815 MB4 GB RAM
ollama run qwen3.5:2b~2.7 GB8 GB RAM
ollama run gemma3:4b~3.3 GB8 GB RAM
ollama run gemma4~9.6 GB16 GB RAM

If a model swaps to disk it still answers, just very slowly — that is the signal to drop a size.

Housekeeping

ollama list
ollama ps
ollama rm gemma3:4b

Lists what you downloaded, shows what is loaded in memory right now, and deletes a model you no longer want.


Next: call your local model from Python and curl or pick a GGUF quantization.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list