Choosing tools

Fine-Tuning vs RAG vs Prompting: How to Choose

3 min read

Almost every “should we fine-tune?” question is really “what is actually broken?” Answer that first and the choice makes itself.

The decision table

PromptingRAGFine-tuning
FixesFormat, tone, reasoning stepsMissing or private factsFormat at scale, niche style
Time to first resultMinutesA dayWeeks
Cost shapePer tokenPer token + index upkeepTraining + hosting + redo on model change
Added latencyNoneOne retrieval hopNone (can shorten prompts)
Goes staleNoNo — reindexYes, as your data moves
DebuggableRead the promptInspect retrieved chunksLargely opaque

Start with prompting

If the output is the wrong shape, tone, or length, that’s a prompt problem. Give it the rules explicitly, add three to five examples of the exact output you want, and separate reasoning from the answer. If a handful of in-prompt examples fixes it, you are done — stop here and don’t build a pipeline.

Reach for RAG when the facts move

The model doesn’t know your docs, your tickets, or last week’s pricing, and no amount of prompting adds them. Retrieval does, and it stays current: change the source document and the next answer changes. It also gives you citations, which is the only cheap way to make an answer checkable.

The tell that you need RAG: the answers are well-written and wrong about your specifics.

Fine-tune for form, not for facts

Fine-tuning shifts behaviour — a house style, a rigid output format, a classification boundary the model keeps missing. It is a bad way to teach facts: they’re baked in, undated, unciteable, and stale the moment the source changes.

It earns its keep in one situation: you already know the right behaviour, you can produce a few hundred to a few thousand clean examples of it, the volume is high, and the few-shot examples you’d otherwise prepend are eating real money on every call. Then you’re trading training cost for a shorter prompt on millions of requests.

Fine-tuning availability differs by provider and by model, and it resets whenever you migrate to a newer base model — check the provider’s current docs before you plan a quarter around it.

Rules of thumb


Next: RAG explained simply · which AI model to use.

Open the full interactive version (with copy buttons) ↗

← All guides