Local LLMs

How to Run Ollama on Windows (Install, GPU, Models)

3 min read

Ollama runs on Windows natively: download OllamaSetup.exe, run it, then ollama run <model>. No administrator rights, no WSL, nothing leaves the machine.

1. Check what Windows needs

The floors come from Ollama’s Windows documentation:

RequirementWhat the docs state
Windows10 22H2 or newer, Home or Pro
NVIDIA driver551.61 or newer
AMDa ROCm v7 / HIP7-capable or Vulkan-capable driver stack
Free diskat least 4 GB for the binaries, plus room for models

2. Install Ollama

irm https://ollama.com/install.ps1 | iex

What it does: the one-liner from the download page; OllamaSetup.exe is the same install with a wizard. It lands in your home directory, so no elevation prompt.

Short on space on C:? The installer takes a directory:

.\OllamaSetup.exe /DIR="D:\ollama"

What it does: puts the binaries where you point it.

3. Pull and run your first model

ollama run gemma4

What it does: downloads the model on first use — several GB — then opens a chat; /bye exits. Any library tag works in its place; check its size against what fits in your RAM or VRAM and the quantization you pick.

4. Verify it worked

ollama ps

What it does: lists the models loaded right now, with the memory each holds. Run it while a reply is still streaming — an idle Ollama has nothing loaded.

curl.exe http://localhost:11434/api/tags

What it does: asks the local server for its models. JSON back means Ollama serves on port 11434 and your code can call that API. Use curl.exe, not curl — PowerShell aliases the bare name to Invoke-WebRequest. Git Bash has the real one.

5. Move the model store off C:

Models land in %HOMEPATH%\.ollama and grow fast. The docs move the store with the OLLAMA_MODELS variable, set through Windows Settings — the same User variable from a terminal:

[Environment]::SetEnvironmentVariable('OLLAMA_MODELS', 'D:\ollama-models', 'User')

What it does: points the store at another drive. Quit Ollama from the tray and start it again; the server reads the variable at startup.

Failure mode: it answers, but at one word a second

That is the CPU doing the work. Two causes worth checking: a graphics driver below the floors above, and a model bigger than your VRAM. Update the driver first; if nothing changes, drop to a smaller tag or a lower quantization.

Ollama’s Windows page documents the native build only. The Linux build inside WSL 2 is a separate install with its own model directory.

Open the full interactive version (with copy buttons) ↗

← All guides