# Call Your Local Ollama Model from curl & Python

> Ollama serves a local API on port 11434. Call /api/chat with curl or Python, or point any OpenAI client at http://localhost:11434/v1/ instead.

- Canonical: https://guides-ai.pages.dev/guides/call-ollama-api/
- Plate 09.03 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 21 Aug 2026 · Updated: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

When Ollama is running it serves an HTTP API at `http://localhost:11434` — no cloud, no key,
no rate limit. Two shapes are available: Ollama's own `/api/*` endpoints, and an
[OpenAI-compatible](/glossary/#openai-compatible-api) surface at `/v1/`.

## 1. Confirm the server and the model

```bash
ollama list
```

What it does: prints the models already on disk. Use one of these names below; if the list is
empty, `ollama pull <model>` first. The desktop app starts the server for you — otherwise run
`ollama serve`.

## 2. curl

```bash
curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{"role": "user", "content": "Explain HTTP 429 in one sentence."}],
  "stream": false,
  "options": {"temperature": 0.2}
}'
```

What it does: one chat request, returning a single JSON object instead of a token stream.
`/api/generate` is the same idea with a plain `prompt` string in place of `messages`.

## 3. Python

```python
import requests

r = requests.post("http://localhost:11434/api/chat", json={
    "model": "gemma4",
    "messages": [{"role": "user", "content": "Give me 3 commit message tips."}],
    "stream": False,
})
print(r.json()["message"]["content"])
```

What it does: the same call from Python; the reply text sits at `message.content`.

## 4. Point an existing OpenAI client at it

```python
from openai import OpenAI

client = OpenAI(
    base_url='http://localhost:11434/v1/',
    api_key='ollama',  # required but ignored
)
resp = client.chat.completions.create(
    model='gemma4',
    messages=[{'role': 'user', 'content': 'Say this is a test'}],
)
print(resp.choices[0].message.content)
```

What it does: reuses code written for OpenAI by changing only the base URL and the model name.
The key is required by the client library and ignored by Ollama. Supported endpoints are
`/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/models` and `/v1/responses`.

## If it fails

`Connection refused` means the server isn't running — start the app or `ollama serve`. A 404
naming the model means it isn't pulled: check `ollama list` against the exact tag, including
the part after the colon. A request that hangs on the first call is usually the model loading
into memory; see [how much RAM a local model needs](/guides/local-llm-ram-vram-requirements/).
Ollama's API isn't strictly versioned but is expected to stay backwards compatible.

Next: [give a local model a baked-in system prompt](/guides/ollama-custom-model-modelfile/) or
[stream responses token by token](/guides/stream-llm-responses/).

Source: [Ollama API documentation](https://docs.ollama.com/api).
