# A Minimal RAG in Python: Folder In, Cited Answer Out

> Chunk a folder of Markdown, embed it with Voyage, retrieve top-k with a numpy dot product, and answer with Claude citing the chunks it used.

- Canonical: https://guides-ai.pages.dev/guides/rag-with-python-minimal/
- Plate 15.18 · Topic: Building with the API (https://guides-ai.pages.dev/topics/api/)
- Published: 06 Sept 2026 · 4 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

Retrieval-augmented generation has a reputation for needing infrastructure. It doesn't, not at the start: the whole pipeline is one embedding call, a dot product, and one message. Here it is over a folder of Markdown notes.

## 1. Install and set keys

```bash
pip install anthropic voyageai numpy
export ANTHROPIC_API_KEY="your-key"
export VOYAGE_API_KEY="your-key"
```

Both SDKs read those variables automatically. Anthropic does not serve an embeddings endpoint, so the retrieval half needs a second provider — Voyage is the one Anthropic's docs point at, and its `input_type` field lets you embed a *question* differently from a *passage*, which is exactly the asymmetry retrieval needs.

## 2–5. The whole thing

```python
import glob, os, anthropic, numpy as np, voyageai

vo, claude = voyageai.Client(), anthropic.Anthropic()

def chunks(folder, size=1200, overlap=200):
    out = []
    for path in glob.glob(os.path.join(folder, "**/*.md"), recursive=True):
        text = open(path, encoding="utf-8").read()
        for i in range(0, len(text), size - overlap):
            piece = text[i:i + size].strip()
            if piece:
                out.append((os.path.basename(path), piece))
    return out

docs = chunks("./notes")
vecs = np.array(vo.embed([c for _, c in docs],
                         model="voyage-4-lite",
                         input_type="document").embeddings)

def ask(question, k=4):
    q = np.array(vo.embed([question], model="voyage-4-lite",
                          input_type="query").embeddings[0])
    scores = vecs @ q / (np.linalg.norm(vecs, axis=1) * np.linalg.norm(q))
    top = np.argsort(-scores)[:k]
    sources = "\n\n".join(f"[{n}] {docs[i][0]}\n{docs[i][1]}"
                          for n, i in enumerate(top, 1))
    msg = claude.messages.create(
        model="claude-opus-5",
        max_tokens=1024,
        system="Answer only from the numbered sources. Cite each claim as [n]. "
               "If the sources do not cover the question, say exactly that.",
        messages=[{"role": "user",
                   "content": f"<sources>\n{sources}\n</sources>\n\n{question}"}],
    )
    return "".join(b.text for b in msg.content if b.type == "text")

print(ask("How do I rotate the deploy key?"))
```

Reads a folder, embeds it once at startup, and answers each question from the four closest chunks. `voyage-4-lite` is the cost-and-latency model; swap in `voyage-4-large` when recall matters more than the bill. The cosine step is one line because `vecs` is a plain matrix — `vecs @ q` scores every chunk at once, and `argsort` picks the winners.

## What to fix first

- **Batching.** The embed endpoint caps how many texts a single call accepts. Slice `docs` into groups before embedding anything bigger than a few dozen files, and concatenate the results.
- **Chunking on character counts** splits mid-table and mid-code-block, which produces confident nonsense downstream. Split on headings first, then by size within a section.
- **The grounding instruction is load-bearing.** Drop "answer only from the numbered sources" and the model quietly fills gaps from memory, with no visible change in tone.
- **Show the sources.** Print the chunk filenames alongside the answer. Most bad answers turn out to be bad retrieval, and you can only see that if you look at what got retrieved.

Move to Chroma or pgvector when the matrix stops fitting in RAM, has to survive a restart, or gets written by more than one process. Until then, an in-memory matrix is faster and has no state to corrupt or migrate.

---

Next: [RAG explained simply](/guides/rag-explained-simply/) · [reduce hallucinations](/guides/reduce-ai-hallucinations/).
