# Fine-Tuning vs RAG vs Prompting: How to Choose

> A decision table plus rules of thumb — which one fixes your problem, and what each costs in money, latency and maintenance.

- Canonical: https://guides-ai.pages.dev/guides/fine-tuning-vs-rag-vs-prompting/
- Plate 11.05 · Topic: Choosing tools (https://guides-ai.pages.dev/topics/choosing/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Almost every "should we fine-tune?" question is really "what is actually broken?" Answer that first and the choice makes itself.

## The decision table

| | Prompting | RAG | Fine-tuning |
| --- | --- | --- | --- |
| **Fixes** | Format, tone, reasoning steps | Missing or private facts | Format at scale, niche style |
| **Time to first result** | Minutes | A day | Weeks |
| **Cost shape** | Per token | Per token + index upkeep | Training + hosting + redo on model change |
| **Added latency** | None | One retrieval hop | None (can shorten prompts) |
| **Goes stale** | No | No — reindex | Yes, as your data moves |
| **Debuggable** | Read the prompt | Inspect retrieved chunks | Largely opaque |

## Start with prompting

If the output is the wrong shape, tone, or length, that's a prompt problem. Give it the rules explicitly, add three to five examples of the exact output you want, and separate reasoning from the answer. If a handful of in-prompt examples fixes it, you are done — stop here and don't build a pipeline.

## Reach for RAG when the facts move

The model doesn't know your docs, your tickets, or last week's pricing, and no amount of prompting adds them. Retrieval does, and it stays current: change the source document and the next answer changes. It also gives you citations, which is the only cheap way to make an answer checkable.

The tell that you need RAG: the answers are well-written and wrong about *your* specifics.

## Fine-tune for form, not for facts

Fine-tuning shifts behaviour — a house style, a rigid output format, a classification boundary the model keeps missing. It is a bad way to teach facts: they're baked in, undated, unciteable, and stale the moment the source changes.

It earns its keep in one situation: you already know the right behaviour, you can produce a few hundred to a few thousand clean examples of it, the volume is high, and the few-shot examples you'd otherwise prepend are eating real money on every call. Then you're trading training cost for a shorter prompt on millions of requests.

Fine-tuning availability differs by provider and by model, and it resets whenever you migrate to a newer base model — check the provider's current docs before you plan a quarter around it.

## Rules of thumb

- Wrong shape → prompt. Wrong facts → RAG. Right facts, wrong shape, high volume → consider fine-tuning.
- Never fine-tune to encode data that changes.
- The combination is normal and often correct: fine-tune the format, retrieve the facts, prompt the task.
- Build an eval *before* you pick, so you can tell whether the change helped.

---

Next: [RAG explained simply](/guides/rag-explained-simply/) · [which AI model to use](/guides/pick-an-ai-model/).
