# How to Run Ollama in Docker with an NVIDIA GPU

> NVIDIA Container Toolkit setup, the official docker run command, pulling a model, keeping models in a volume, and AMD or CPU fallbacks.

- Canonical: https://guides-ai.pages.dev/guides/ollama-in-docker-with-gpu/
- Plate 09.09 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Docker keeps Ollama and its multi-gigabyte model cache out of your system install, which makes it easy to move, upgrade or throw away. The only real work is giving the container access to the GPU.

## 1. Install the NVIDIA Container Toolkit

Linux hosts need it — Docker cannot see your GPU without it. Add NVIDIA's package repository (the repo lines are on NVIDIA's Container Toolkit installation page), then:

```bash
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
```

Installs the toolkit, registers it as a Docker runtime, and restarts the daemon so the change takes effect. Verify before going further:

```bash
docker run --rm --gpus=all ubuntu nvidia-smi
```

Prints your GPU table from inside a container. If this fails, Ollama will fail too — fix it here.

**Windows and macOS.** On Windows, Docker Desktop with the WSL 2 backend passes an NVIDIA GPU through with a current driver; there is no separate toolkit to install. On macOS, Docker has no GPU passthrough at all — the container will be CPU-only, so run the native Ollama app instead.

## 2. Start Ollama

```bash
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
```

The official command. Runs detached, hands the container every GPU, publishes the API on port 11434, and stores everything in a named volume called `ollama` mounted at `/root/.ollama`.

That volume is the important part: models live there, not in the container, so `docker rm` never costs you a re-download. Prefer a visible folder? Swap in `-v ~/.ollama:/root/.ollama`.

## 3. Pull and run a model

```bash
docker exec -it ollama ollama run llama3.2
```

Downloads the model on first use and drops you into a chat prompt inside the container. From the host, the API works exactly as a native install does:

```bash
curl http://localhost:11434/api/tags
```

Lists the installed models — proof the port mapping works.

## 4. Confirm it is on the GPU

```bash
docker exec ollama ollama ps
```

Shows loaded models and a `PROCESSOR` column. `100% GPU` means the weights fit on the card; anything mentioning CPU means a partial or full CPU fallback — usually a model too large for your VRAM.

## 5. AMD, CPU, upgrades

```bash
docker run -d --device /dev/kfd --device /dev/dri -v ollama:/root/.ollama \
  -p 11434:11434 --name ollama ollama/ollama:rocm
```

The ROCm image for AMD cards — note the different tag and the two device passthroughs instead of `--gpus`. For CPU-only, use the plain command from step 2 with `--gpus=all` removed.

To upgrade, `docker rm -f ollama && docker pull ollama/ollama`, then re-run your `docker run`. The volume survives, so your models do.

---

Next: [add a chat interface with Open WebUI](/guides/open-webui-with-ollama/) or [call the API from your own code](/guides/call-ollama-api/).
