Docker keeps Ollama and its multi-gigabyte model cache out of your system install, which makes it easy to move, upgrade or throw away. The only real work is giving the container access to the GPU.
1. Install the NVIDIA Container Toolkit
Linux hosts need it — Docker cannot see your GPU without it. Add NVIDIA’s package repository (the repo lines are on NVIDIA’s Container Toolkit installation page), then:
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Installs the toolkit, registers it as a Docker runtime, and restarts the daemon so the change takes effect. Verify before going further:
docker run --rm --gpus=all ubuntu nvidia-smi
Prints your GPU table from inside a container. If this fails, Ollama will fail too — fix it here.
Windows and macOS. On Windows, Docker Desktop with the WSL 2 backend passes an NVIDIA GPU through with a current driver; there is no separate toolkit to install. On macOS, Docker has no GPU passthrough at all — the container will be CPU-only, so run the native Ollama app instead.
2. Start Ollama
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
The official command. Runs detached, hands the container every GPU, publishes the API on port 11434, and stores everything in a named volume called ollama mounted at /root/.ollama.
That volume is the important part: models live there, not in the container, so docker rm never costs you a re-download. Prefer a visible folder? Swap in -v ~/.ollama:/root/.ollama.
3. Pull and run a model
docker exec -it ollama ollama run llama3.2
Downloads the model on first use and drops you into a chat prompt inside the container. From the host, the API works exactly as a native install does:
curl http://localhost:11434/api/tags
Lists the installed models — proof the port mapping works.
4. Confirm it is on the GPU
docker exec ollama ollama ps
Shows loaded models and a PROCESSOR column. 100% GPU means the weights fit on the card; anything mentioning CPU means a partial or full CPU fallback — usually a model too large for your VRAM.
5. AMD, CPU, upgrades
docker run -d --device /dev/kfd --device /dev/dri -v ollama:/root/.ollama \
-p 11434:11434 --name ollama ollama/ollama:rocm
The ROCm image for AMD cards — note the different tag and the two device passthroughs instead of --gpus. For CPU-only, use the plain command from step 2 with --gpus=all removed.
To upgrade, docker rm -f ollama && docker pull ollama/ollama, then re-run your docker run. The volume survives, so your models do.
Next: add a chat interface with Open WebUI or call the API from your own code.