A 90-page contract will not fit in a small local model’s context, and you probably do not want to upload it anyway. The fix is the standard map-reduce shape: summarize each chunk, then summarize the summaries. It costs nothing, runs offline, and the intermediate summaries are useful on their own — they are a section-by-section outline of the document.
1. Install
pip install pypdf requests
ollama pull llama3.2
pypdf reads the PDF, requests talks to Ollama. Check the server is up before running anything: curl http://localhost:11434/api/tags should return JSON.
2. The whole script
import requests
from pypdf import PdfReader
MODEL, URL = "llama3.2", "http://localhost:11434/api/generate"
CHUNK = 6000 # characters, roughly 1,500 tokens
def ask(prompt: str) -> str:
r = requests.post(URL, json={"model": MODEL, "prompt": prompt, "stream": False})
r.raise_for_status()
return r.json()["response"].strip()
pages = PdfReader("report.pdf").pages
text = "\n".join((p.extract_text() or "") for p in pages)
print(f"{len(pages)} pages, {len(text)} characters")
chunks = [text[i:i + CHUNK] for i in range(0, len(text), CHUNK)]
parts = []
for n, chunk in enumerate(chunks, 1):
print(f"summarizing {n}/{len(chunks)}")
parts.append(ask(
"Summarize this section of a document in 3 bullet points. "
"Use only what is written below; do not add background knowledge.\n\n" + chunk
))
summary = ask(
"These are section summaries of one document, in order. Merge them into a single "
"200-word summary with no repetition and no preamble:\n\n" + "\n\n".join(parts)
)
print("\n" + summary)
Reads every page, splits the text into 6,000-character blocks, sends each to your local model, then sends the collected bullet points back for one final pass. Save it as summarize.py and run python summarize.py. Progress prints as it goes, because a long PDF on a small machine takes minutes.
3. When it misbehaves
Empty output, no error. extract_text() returned nothing because the PDF is scanned images. The page count will print fine while the character count is near zero. Run it through OCR first (ocrmypdf in.pdf out.pdf) and try again.
Summaries that ignore the end of a chunk. The chunk is longer than the context window the model was loaded with. Lower CHUNK, or raise the window explicitly:
json={"model": MODEL, "prompt": prompt, "stream": False,
"options": {"num_ctx": 8192, "temperature": 0.2}}
Gives the model a larger context and steadier, less inventive output.
Too slow. Use a smaller model for the per-chunk pass — the merge step is the one that needs judgment. parts is just a list of strings, so you can print it, cache it, or feed it to a bigger model at the end.
Next: call the Ollama API properly or do the same job in ChatGPT.