Local LLMs

Call Your Local Ollama Model from curl & Python

3 min read

When Ollama is running it serves an HTTP API at http://localhost:11434 — no cloud, no key, no rate limit. Two shapes are available: Ollama’s own /api/* endpoints, and an OpenAI-compatible surface at /v1/.

1. Confirm the server and the model

ollama list

What it does: prints the models already on disk. Use one of these names below; if the list is empty, ollama pull <model> first. The desktop app starts the server for you — otherwise run ollama serve.

2. curl

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{"role": "user", "content": "Explain HTTP 429 in one sentence."}],
  "stream": false,
  "options": {"temperature": 0.2}
}'

What it does: one chat request, returning a single JSON object instead of a token stream. /api/generate is the same idea with a plain prompt string in place of messages.

3. Python

import requests

r = requests.post("http://localhost:11434/api/chat", json={
    "model": "gemma4",
    "messages": [{"role": "user", "content": "Give me 3 commit message tips."}],
    "stream": False,
})
print(r.json()["message"]["content"])

What it does: the same call from Python; the reply text sits at message.content.

4. Point an existing OpenAI client at it

from openai import OpenAI

client = OpenAI(
    base_url='http://localhost:11434/v1/',
    api_key='ollama',  # required but ignored
)
resp = client.chat.completions.create(
    model='gemma4',
    messages=[{'role': 'user', 'content': 'Say this is a test'}],
)
print(resp.choices[0].message.content)

What it does: reuses code written for OpenAI by changing only the base URL and the model name. The key is required by the client library and ignored by Ollama. Supported endpoints are /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models and /v1/responses.

If it fails

Connection refused means the server isn’t running — start the app or ollama serve. A 404 naming the model means it isn’t pulled: check ollama list against the exact tag, including the part after the colon. A request that hangs on the first call is usually the model loading into memory; see how much RAM a local model needs. Ollama’s API isn’t strictly versioned but is expected to stay backwards compatible.

Next: give a local model a baked-in system prompt or stream responses token by token.

Source: Ollama API documentation.

Open the full interactive version (with copy buttons) ↗

← All guides