Call Your Local Ollama Model from curl & Python
Ollama serves a local API on port 11434. Call /api/chat with curl or Python, or point any OpenAI client at http://localhost:11434/v1/ instead.
Step 4 of 5 · Run models locally
On this page5 sections
When Ollama is running it serves an HTTP API at http://localhost:11434 — no cloud, no key,
no rate limit. Two shapes are available: Ollama’s own /api/* endpoints, and an
OpenAI-compatible surface at /v1/.
1. Confirm the server and the model
ollama list
What it does: prints the models already on disk. Use one of these names below; if the list is
empty, ollama pull <model> first. The desktop app starts the server for you — otherwise run
ollama serve.
2. curl
curl http://localhost:11434/api/chat -d '{
"model": "gemma4",
"messages": [{"role": "user", "content": "Explain HTTP 429 in one sentence."}],
"stream": false,
"options": {"temperature": 0.2}
}'
What it does: one chat request, returning a single JSON object instead of a token stream.
/api/generate is the same idea with a plain prompt string in place of messages.
3. Python
import requests
r = requests.post("http://localhost:11434/api/chat", json={
"model": "gemma4",
"messages": [{"role": "user", "content": "Give me 3 commit message tips."}],
"stream": False,
})
print(r.json()["message"]["content"])
What it does: the same call from Python; the reply text sits at message.content.
4. Point an existing OpenAI client at it
from openai import OpenAI
client = OpenAI(
base_url='http://localhost:11434/v1/',
api_key='ollama', # required but ignored
)
resp = client.chat.completions.create(
model='gemma4',
messages=[{'role': 'user', 'content': 'Say this is a test'}],
)
print(resp.choices[0].message.content)
What it does: reuses code written for OpenAI by changing only the base URL and the model name.
The key is required by the client library and ignored by Ollama. Supported endpoints are
/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models and /v1/responses.
If it fails
Connection refused means the server isn’t running — start the app or ollama serve. A 404
naming the model means it isn’t pulled: check ollama list against the exact tag, including
the part after the colon. A request that hangs on the first call is usually the model loading
into memory; see how much RAM a local model needs.
Ollama’s API isn’t strictly versioned but is expected to stay backwards compatible.
Next: give a local model a baked-in system prompt or stream responses token by token.
Source: Ollama API documentation.