# How to Run a Local LLM with Ollama on macOS/Linux

> Install Ollama with one command, pull gemma3:4b, chat offline in the terminal, and keep the daemon running as a background service on port 11434.

- Canonical: https://guides-ai.pages.dev/guides/run-local-llm-ollama-macos-linux/
- Plate 09.11 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Ollama runs open models on your own machine. Same install command on macOS and Linux, no API key.

## 1. Install

```bash
curl -fsSL https://ollama.com/install.sh | sh
```

On macOS this installs `Ollama.app` into `/Applications` and links the `ollama` command into `/usr/local/bin`. On Linux it installs the binary and writes a systemd unit at `/etc/systemd/system/ollama.service`.

Check it worked:

```bash
ollama -v
```

Prints the installed version — if the command is not found, open a new terminal so your `PATH` refreshes.

## 2. Pull a small model

```bash
ollama pull gemma3:4b
```

Downloads a ~3.3 GB 4B model. Start here rather than with the default `gemma4`, which is a ~9.6 GB download.

## 3. Chat

```bash
ollama run gemma3:4b
```

Loads the model and drops you into a prompt. Type `/bye` to leave the chat and unload it.

## 4. Keep it running as a service

On **Linux**, the installer already enabled the service. To check, control, and watch it:

```bash
systemctl status ollama
sudo systemctl enable --now ollama
journalctl -u ollama -f
```

Shows whether the daemon is up, makes it start at boot, and tails its log. The API is then always available on `http://localhost:11434`.

On **macOS**, the app starts at login and does the same job. If you would rather run it manually in a terminal:

```bash
ollama serve
```

Runs the daemon in the foreground; leave that window open while you use the API.

## Choosing a size for your RAM

You need roughly the download size in free RAM (or VRAM), plus 1–2 GB of headroom for context:

| Command | Download | Comfortable on |
| --- | --- | --- |
| `ollama run gemma3:1b` | ~815 MB | 4 GB RAM |
| `ollama run qwen3.5:2b` | ~2.7 GB | 8 GB RAM |
| `ollama run gemma3:4b` | ~3.3 GB | 8 GB RAM |
| `ollama run gemma4` | ~9.6 GB | 16 GB RAM |

If a model swaps to disk it still answers, just very slowly — that is the signal to drop a size.

## Housekeeping

```bash
ollama list
ollama ps
ollama rm gemma3:4b
```

Lists what you downloaded, shows what is loaded in memory right now, and deletes a model you no longer want.

---

Next: [call your local model from Python and curl](/guides/call-ollama-api/) or [pick a GGUF quantization](/guides/choose-llm-quantization/).
