§15.17

RAG Without a Vector Database: OpenAI File Search

Upload files to an OpenAI vector store, attach the file_search tool to a Responses API call, and get cited answers — no embeddings code, no database.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Building with the API Markdown

On this page5 sections
  1. 1. Create a store and upload files
  2. 2. Ask a question
  3. 3. See what it retrieved
  4. 4. Filter by metadata
  5. When to build your own instead

If you want “answer questions from these 40 PDFs” without owning an embeddings pipeline, OpenAI’s file search does the chunking, embedding and retrieval on their side. You upload files into a vector store, hand the store’s id to the file_search tool, and the model retrieves before it answers.

1. Create a store and upload files

from openai import OpenAI
client = OpenAI()  # reads OPENAI_API_KEY

store = client.vector_stores.create(name="knowledge_base")

with open("handbook.pdf", "rb") as f:
    uploaded = client.files.create(file=f, purpose="assistants")

client.vector_stores.files.create(vector_store_id=store.id, file_id=uploaded.id)
print(store.id)

Creates the store, uploads one file with purpose="assistants", and attaches it. Repeat the upload + attach for each file; indexing runs in the background, so give a large batch a moment before the first query. Supported formats include .pdf, .docx, .txt, .md, .json and common code files.

2. Ask a question

response = client.responses.create(
    model="gpt-6-astra",
    input="What is the parental leave policy? Cite the section.",
    tools=[{"type": "file_search", "vector_store_ids": [store.id]}],
)
print(response.output_text)

The model decides when to search; the retrieved chunks go into its context and the answer comes back as normal text. vector_store_ids is a list — you can search several stores in one call.

3. See what it retrieved

response = client.responses.create(
    model="gpt-6-astra",
    input="Summarize the expense rules.",
    tools=[{"type": "file_search", "vector_store_ids": [store.id], "max_num_results": 4}],
    include=["file_search_call.results"],
)
for item in response.output:
    if item.type == "file_search_call":
        print(item.results)          # the chunks that were used
    elif item.type == "message":
        for part in item.content:
            for a in getattr(part, "annotations", []):
                print(a.filename, a.file_id)   # file citations

max_num_results caps how many chunks are injected (fewer = cheaper, sometimes worse); include=["file_search_call.results"] returns the raw hits so you can show sources. Citations arrive as annotations on the message content, with filename and file_id.

4. Filter by metadata

tools=[{
    "type": "file_search",
    "vector_store_ids": [store.id],
    "filters": {"type": "in", "key": "category", "value": ["policy"]},
}]

Works when you attached files with attributes (e.g. {"category": "policy"}) — handy for one store that serves several teams.

When to build your own instead

You control nothing about chunking or the embedding model here, storage is billed by the GB, and your data lives in OpenAI’s store. For a private corpus, an on-prem requirement, or when you want to tune retrieval, the 40-line RAG in Python is the other path.


Next: a minimal RAG in Python · embeddings explained with code.

← All Building with the API plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list