§11.01

Which AI Model Should You Use?

Pick an AI model by job — coding, images, very long inputs, or private local inference — with the current tiers and the context-window myth corrected.

published 10 Jun 2026 updated 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Choosing tools Markdown

Step 6 of 6 · Get running

On this page4 sections
  1. By job
  2. The context-window rule of thumb is out of date
  3. Tiers matter more than brands
  4. When local wins

Start at the mid tier of whichever provider you already pay for, and move only when a real task fails. Every vendor’s second-from-top model is close enough to the flagship for ordinary work at a fraction of the price, and swapping brands rarely fixes what a better prompt would.

By job

JobReach for
Coding, long documents, careful writingClaude — Anthropic’s own guidance is to start with Opus 5
Images, voice, one broad ecosystemChatGPT / OpenAI
Google Workspace and Cloud integrationGemini
Private data, offline, no per-token costOpen weights via Ollama

The context-window rule of thumb is out of date

“Use Gemini when the input is huge” was true for a while. It isn’t now: Claude Opus 5 and Sonnet 5 carry 1M-token windows, and OpenAI’s GPT-6 and GPT-5.6 models carry about 1.05M. A million tokens is roughly 555,000 words — a long book. Pick on integration and price instead, and remember that a giant context window is something you pay for by the token.

Tiers matter more than brands

Each vendor ships a fast cheap model, a balanced one, and a slow expensive one. The balanced tier — Claude Sonnet 5, OpenAI’s GPT-5.6 Terra — handles most production work. Small models like Claude Haiku 4.5 are for classification, extraction and routing, where you make thousands of calls and none needs deep reasoning. The full price comparison is in OpenAI API vs Claude API.

When local wins

If the data cannot leave your machine, or you need unlimited volume with no per-token bill, run open weights yourself. An 8B model handles summarising, extraction and drafting; check it fits with RAM and VRAM requirements, then:

ollama run llama3.2 "Summarise this in three bullets: <paste text>"

Runs entirely offline after the first pull — no key, no quota, no data leaving the box. Start here: run a local LLM with Ollama.

Then re-check every few months. Tiers get re-cut and repriced two or three times a year, so last winter’s choice often pays yesterday’s rate for yesterday’s model — re-read the Claude models overview before pinning anything in production.

← All Choosing tools plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list