§11.05

Fine-Tuning vs RAG vs Prompting: How to Choose

A decision table plus rules of thumb — which one fixes your problem, and what each costs in money, latency and maintenance.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Choosing tools Markdown

On this page5 sections
  1. The decision table
  2. Start with prompting
  3. Reach for RAG when the facts move
  4. Fine-tune for form, not for facts
  5. Rules of thumb

Almost every “should we fine-tune?” question is really “what is actually broken?” Answer that first and the choice makes itself.

The decision table

PromptingRAGFine-tuning
FixesFormat, tone, reasoning stepsMissing or private factsFormat at scale, niche style
Time to first resultMinutesA dayWeeks
Cost shapePer tokenPer token + index upkeepTraining + hosting + redo on model change
Added latencyNoneOne retrieval hopNone (can shorten prompts)
Goes staleNoNo — reindexYes, as your data moves
DebuggableRead the promptInspect retrieved chunksLargely opaque

Start with prompting

If the output is the wrong shape, tone, or length, that’s a prompt problem. Give it the rules explicitly, add three to five examples of the exact output you want, and separate reasoning from the answer. If a handful of in-prompt examples fixes it, you are done — stop here and don’t build a pipeline.

Reach for RAG when the facts move

The model doesn’t know your docs, your tickets, or last week’s pricing, and no amount of prompting adds them. Retrieval does, and it stays current: change the source document and the next answer changes. It also gives you citations, which is the only cheap way to make an answer checkable.

The tell that you need RAG: the answers are well-written and wrong about your specifics.

Fine-tune for form, not for facts

Fine-tuning shifts behaviour — a house style, a rigid output format, a classification boundary the model keeps missing. It is a bad way to teach facts: they’re baked in, undated, unciteable, and stale the moment the source changes.

It earns its keep in one situation: you already know the right behaviour, you can produce a few hundred to a few thousand clean examples of it, the volume is high, and the few-shot examples you’d otherwise prepend are eating real money on every call. Then you’re trading training cost for a shorter prompt on millions of requests.

Fine-tuning availability differs by provider and by model, and it resets whenever you migrate to a newer base model — check the provider’s current docs before you plan a quarter around it.

Rules of thumb

  • Wrong shape → prompt. Wrong facts → RAG. Right facts, wrong shape, high volume → consider fine-tuning.
  • Never fine-tune to encode data that changes.
  • The combination is normal and often correct: fine-tune the format, retrieve the facts, prompt the task.
  • Build an eval before you pick, so you can tell whether the change helped.

Next: RAG explained simply · which AI model to use.

← All Choosing tools plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list