RAG Without a Vector Database: OpenAI File Search
Upload files to an OpenAI vector store, attach the file_search tool to a Responses API call, and get cited answers — no embeddings code, no database.
On this page5 sections
If you want “answer questions from these 40 PDFs” without owning an embeddings pipeline, OpenAI’s file search does the chunking, embedding and retrieval on their side. You upload files into a vector store, hand the store’s id to the file_search tool, and the model retrieves before it answers.
1. Create a store and upload files
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY
store = client.vector_stores.create(name="knowledge_base")
with open("handbook.pdf", "rb") as f:
uploaded = client.files.create(file=f, purpose="assistants")
client.vector_stores.files.create(vector_store_id=store.id, file_id=uploaded.id)
print(store.id)
Creates the store, uploads one file with purpose="assistants", and attaches it. Repeat the upload + attach for each file; indexing runs in the background, so give a large batch a moment before the first query. Supported formats include .pdf, .docx, .txt, .md, .json and common code files.
2. Ask a question
response = client.responses.create(
model="gpt-6-astra",
input="What is the parental leave policy? Cite the section.",
tools=[{"type": "file_search", "vector_store_ids": [store.id]}],
)
print(response.output_text)
The model decides when to search; the retrieved chunks go into its context and the answer comes back as normal text. vector_store_ids is a list — you can search several stores in one call.
3. See what it retrieved
response = client.responses.create(
model="gpt-6-astra",
input="Summarize the expense rules.",
tools=[{"type": "file_search", "vector_store_ids": [store.id], "max_num_results": 4}],
include=["file_search_call.results"],
)
for item in response.output:
if item.type == "file_search_call":
print(item.results) # the chunks that were used
elif item.type == "message":
for part in item.content:
for a in getattr(part, "annotations", []):
print(a.filename, a.file_id) # file citations
max_num_results caps how many chunks are injected (fewer = cheaper, sometimes worse); include=["file_search_call.results"] returns the raw hits so you can show sources. Citations arrive as annotations on the message content, with filename and file_id.
4. Filter by metadata
tools=[{
"type": "file_search",
"vector_store_ids": [store.id],
"filters": {"type": "in", "key": "category", "value": ["policy"]},
}]
Works when you attached files with attributes (e.g. {"category": "policy"}) — handy for one store that serves several teams.
When to build your own instead
You control nothing about chunking or the embedding model here, storage is billed by the GB, and your data lives in OpenAI’s store. For a private corpus, an on-prem requirement, or when you want to tune retrieval, the 40-line RAG in Python is the other path.
Next: a minimal RAG in Python · embeddings explained with code.