Start at the mid tier of whichever provider you already pay for, and move only when a real task fails. Every vendor’s second-from-top model is close enough to the flagship for ordinary work at a fraction of the price, and swapping brands rarely fixes what a better prompt would.
By job
| Job | Reach for |
|---|---|
| Coding, long documents, careful writing | Claude — Anthropic’s own guidance is to start with Opus 5 |
| Images, voice, one broad ecosystem | ChatGPT / OpenAI |
| Google Workspace and Cloud integration | Gemini |
| Private data, offline, no per-token cost | Open weights via Ollama |
The context-window rule of thumb is out of date
“Use Gemini when the input is huge” was true for a while. It isn’t now: Claude Opus 5 and Sonnet 5 carry 1M-token windows, and OpenAI’s GPT-6 and GPT-5.6 models carry about 1.05M. A million tokens is roughly 555,000 words — a long book. Pick on integration and price instead, and remember that a giant context window is something you pay for by the token.
Tiers matter more than brands
Each vendor ships a fast cheap model, a balanced one, and a slow expensive one. The balanced tier — Claude Sonnet 5, OpenAI’s GPT-5.6 Terra — handles most production work. Small models like Claude Haiku 4.5 are for classification, extraction and routing, where you make thousands of calls and none needs deep reasoning. The full price comparison is in OpenAI API vs Claude API.
When local wins
If the data cannot leave your machine, or you need unlimited volume with no per-token bill, run open weights yourself. An 8B model handles summarising, extraction and drafting; check it fits with RAM and VRAM requirements, then:
ollama run llama3.2 "Summarise this in three bullets: <paste text>"
Runs entirely offline after the first pull — no key, no quota, no data leaving the box. Start here: run a local LLM with Ollama.
Then re-check every few months. Tiers get re-cut and repriced two or three times a year, so last winter’s choice often pays yesterday’s rate for yesterday’s model — re-read the Claude models overview before pinning anything in production.