# RAG Explained Simply (and When You Need It)

> RAG retrieves relevant chunks, then generates. With 1M-token context windows and hosted file_search tools, here is when it still beats pasting the file.

- Canonical: https://guides-ai.pages.dev/guides/rag-explained-simply/
- Plate 10.05 · Topic: Prompting (https://guides-ai.pages.dev/topics/prompting/)
- Published: 02 Sept 2026 · Updated: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

**[RAG](/glossary/#rag) means: search your own data first, then put the results in the prompt.** The model answers from text you handed it rather than from memory. No retraining, no fine-tuning — retrieve, then generate.

## The flow

```text
1. Split documents into chunks and embed them once.
2. Embed the question, find the nearest chunks.
3. Put those chunks in the prompt as context.
4. The model answers from them, and can cite them.
```

The whole pattern in four lines. Steps 1 and 2 are ordinary search; only step 3 involves the model. See [embeddings with code](/guides/embeddings-explained-with-code/) for what the vectors actually are.

## What changed: context got huge

The old rule was "use RAG when the data doesn't fit." That threshold moved. Current flagship models from Anthropic and OpenAI carry roughly a million tokens of [context window](/glossary/#context-window) — several hundred thousand words. A whole handbook now fits in one prompt.

So RAG is no longer about capacity. It earns its keep when you have far more data than any window holds, when the data changes faster than you can re-paste it, when you need a citation back to a source, or when you refuse to pay for 900k tokens to answer one question. Note that very long prompts can also cost more per token: OpenAI prices requests above 272k input tokens at double the input rate.

## RAG vs fine-tuning

Fine-tuning teaches style, format and behaviour; it does not reliably teach facts, and it goes stale the day your documents change. RAG keeps facts external and swappable — edit the source, no retraining. For "answer questions about our docs," it is almost always RAG. The longer comparison: [fine-tuning vs RAG vs prompting](/guides/fine-tuning-vs-rag-vs-prompting/).

## Don't build it first

Hosted tools do the chunking, embedding, keyword-plus-semantic search and reranking for you:

```python
from openai import OpenAI
client = OpenAI()

store = client.vector_stores.create(name="knowledge_base")
with open("handbook.pdf", "rb") as f:
    up = client.files.create(file=f, purpose="assistants")
client.vector_stores.files.create(vector_store_id=store.id, file_id=up.id)

r = client.responses.create(
    model="gpt-6-astra",
    input="What is our refund window?",
    tools=[{"type": "file_search", "vector_store_ids": [store.id]}],
)
print(r.output_text)
```

Uploads one PDF and answers from it — a working RAG pipeline with no vector database to run. Details: [the file search guide](https://developers.openai.com/api/docs/guides/tools-file-search). Build it yourself only when you need control over chunking or hosting: [RAG with Python, minimal](/guides/rag-with-python-minimal/).
