# RAG Without a Vector Database: OpenAI File Search

> Upload files to an OpenAI vector store, attach the file_search tool to a Responses API call, and get cited answers — no embeddings code, no database.

- Canonical: https://guides-ai.pages.dev/guides/openai-file-search-rag/
- Plate 15.17 · Topic: Building with the API (https://guides-ai.pages.dev/topics/api/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

If you want "answer questions from these 40 PDFs" without owning an embeddings pipeline, OpenAI's **file search** does the chunking, embedding and retrieval on their side. You upload files into a vector store, hand the store's id to the `file_search` tool, and the model retrieves before it answers.

## 1. Create a store and upload files

```python
from openai import OpenAI
client = OpenAI()  # reads OPENAI_API_KEY

store = client.vector_stores.create(name="knowledge_base")

with open("handbook.pdf", "rb") as f:
    uploaded = client.files.create(file=f, purpose="assistants")

client.vector_stores.files.create(vector_store_id=store.id, file_id=uploaded.id)
print(store.id)
```

Creates the store, uploads one file with `purpose="assistants"`, and attaches it. Repeat the upload + attach for each file; indexing runs in the background, so give a large batch a moment before the first query. Supported formats include `.pdf`, `.docx`, `.txt`, `.md`, `.json` and common code files.

## 2. Ask a question

```python
response = client.responses.create(
    model="gpt-6-astra",
    input="What is the parental leave policy? Cite the section.",
    tools=[{"type": "file_search", "vector_store_ids": [store.id]}],
)
print(response.output_text)
```

The model decides when to search; the retrieved chunks go into its context and the answer comes back as normal text. `vector_store_ids` is a list — you can search several stores in one call.

## 3. See what it retrieved

```python
response = client.responses.create(
    model="gpt-6-astra",
    input="Summarize the expense rules.",
    tools=[{"type": "file_search", "vector_store_ids": [store.id], "max_num_results": 4}],
    include=["file_search_call.results"],
)
for item in response.output:
    if item.type == "file_search_call":
        print(item.results)          # the chunks that were used
    elif item.type == "message":
        for part in item.content:
            for a in getattr(part, "annotations", []):
                print(a.filename, a.file_id)   # file citations
```

`max_num_results` caps how many chunks are injected (fewer = cheaper, sometimes worse); `include=["file_search_call.results"]` returns the raw hits so you can show sources. Citations arrive as `annotations` on the message content, with `filename` and `file_id`.

## 4. Filter by metadata

```python
tools=[{
    "type": "file_search",
    "vector_store_ids": [store.id],
    "filters": {"type": "in", "key": "category", "value": ["policy"]},
}]
```

Works when you attached files with `attributes` (e.g. `{"category": "policy"}`) — handy for one store that serves several teams.

## When to build your own instead

You control nothing about chunking or the embedding model here, storage is billed by the GB, and your data lives in OpenAI's store. For a private corpus, an on-prem requirement, or when you want to tune retrieval, the [40-line RAG in Python](/guides/rag-with-python-minimal/) is the other path.

---

Next: [a minimal RAG in Python](/guides/rag-with-python-minimal/) · [embeddings explained with code](/guides/embeddings-explained-with-code/).
