§13.03

How to Summarize a PDF Locally with Ollama

Extract PDF text with pypdf, chunk it, summarize each chunk through the local Ollama API, and merge — about 30 lines, nothing uploaded.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Everyday tasks Markdown

On this page3 sections
  1. 1. Install
  2. 2. The whole script
  3. 3. When it misbehaves

A 90-page contract will not fit in a small local model’s context, and you probably do not want to upload it anyway. The fix is the standard map-reduce shape: summarize each chunk, then summarize the summaries. It costs nothing, runs offline, and the intermediate summaries are useful on their own — they are a section-by-section outline of the document.

1. Install

pip install pypdf requests
ollama pull llama3.2

pypdf reads the PDF, requests talks to Ollama. Check the server is up before running anything: curl http://localhost:11434/api/tags should return JSON.

2. The whole script

import requests
from pypdf import PdfReader

MODEL, URL = "llama3.2", "http://localhost:11434/api/generate"
CHUNK = 6000  # characters, roughly 1,500 tokens

def ask(prompt: str) -> str:
    r = requests.post(URL, json={"model": MODEL, "prompt": prompt, "stream": False})
    r.raise_for_status()
    return r.json()["response"].strip()

pages = PdfReader("report.pdf").pages
text = "\n".join((p.extract_text() or "") for p in pages)
print(f"{len(pages)} pages, {len(text)} characters")

chunks = [text[i:i + CHUNK] for i in range(0, len(text), CHUNK)]
parts = []
for n, chunk in enumerate(chunks, 1):
    print(f"summarizing {n}/{len(chunks)}")
    parts.append(ask(
        "Summarize this section of a document in 3 bullet points. "
        "Use only what is written below; do not add background knowledge.\n\n" + chunk
    ))

summary = ask(
    "These are section summaries of one document, in order. Merge them into a single "
    "200-word summary with no repetition and no preamble:\n\n" + "\n\n".join(parts)
)
print("\n" + summary)

Reads every page, splits the text into 6,000-character blocks, sends each to your local model, then sends the collected bullet points back for one final pass. Save it as summarize.py and run python summarize.py. Progress prints as it goes, because a long PDF on a small machine takes minutes.

3. When it misbehaves

Empty output, no error. extract_text() returned nothing because the PDF is scanned images. The page count will print fine while the character count is near zero. Run it through OCR first (ocrmypdf in.pdf out.pdf) and try again.

Summaries that ignore the end of a chunk. The chunk is longer than the context window the model was loaded with. Lower CHUNK, or raise the window explicitly:

json={"model": MODEL, "prompt": prompt, "stream": False,
      "options": {"num_ctx": 8192, "temperature": 0.2}}

Gives the model a larger context and steadier, less inventive output.

Too slow. Use a smaller model for the per-chunk pass — the merge step is the one that needs judgment. parts is just a list of strings, so you can print it, cache it, or feed it to a bigger model at the end.


Next: call the Ollama API properly or do the same job in ChatGPT.

← All Everyday tasks plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list