§09.03

Call Your Local Ollama Model from curl & Python

Ollama serves a local API on port 11434. Call /api/chat with curl or Python, or point any OpenAI client at http://localhost:11434/v1/ instead.

published 21 Aug 2026 updated 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Local LLMs Markdown

Step 4 of 5 · Run models locally

On this page5 sections
  1. 1. Confirm the server and the model
  2. 2. curl
  3. 3. Python
  4. 4. Point an existing OpenAI client at it
  5. If it fails

When Ollama is running it serves an HTTP API at http://localhost:11434 — no cloud, no key, no rate limit. Two shapes are available: Ollama’s own /api/* endpoints, and an OpenAI-compatible surface at /v1/.

1. Confirm the server and the model

ollama list

What it does: prints the models already on disk. Use one of these names below; if the list is empty, ollama pull <model> first. The desktop app starts the server for you — otherwise run ollama serve.

2. curl

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{"role": "user", "content": "Explain HTTP 429 in one sentence."}],
  "stream": false,
  "options": {"temperature": 0.2}
}'

What it does: one chat request, returning a single JSON object instead of a token stream. /api/generate is the same idea with a plain prompt string in place of messages.

3. Python

import requests

r = requests.post("http://localhost:11434/api/chat", json={
    "model": "gemma4",
    "messages": [{"role": "user", "content": "Give me 3 commit message tips."}],
    "stream": False,
})
print(r.json()["message"]["content"])

What it does: the same call from Python; the reply text sits at message.content.

4. Point an existing OpenAI client at it

from openai import OpenAI

client = OpenAI(
    base_url='http://localhost:11434/v1/',
    api_key='ollama',  # required but ignored
)
resp = client.chat.completions.create(
    model='gemma4',
    messages=[{'role': 'user', 'content': 'Say this is a test'}],
)
print(resp.choices[0].message.content)

What it does: reuses code written for OpenAI by changing only the base URL and the model name. The key is required by the client library and ignored by Ollama. Supported endpoints are /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models and /v1/responses.

If it fails

Connection refused means the server isn’t running — start the app or ollama serve. A 404 naming the model means it isn’t pulled: check ollama list against the exact tag, including the part after the colon. A request that hangs on the first call is usually the model loading into memory; see how much RAM a local model needs. Ollama’s API isn’t strictly versioned but is expected to stay backwards compatible.

Next: give a local model a baked-in system prompt or stream responses token by token.

Source: Ollama API documentation.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list