Skip to content

Getting started

Install

pip install pylopdf

For rendering Japanese PDFs without embedded fonts, or auto-subsetting a JP font for Japanese/Han insert_text and insert_textbox, install the bundled Noto Sans/Serif JP package:

pip install pylopdf[cjk]

Release CI installs every native platform wheel on an architecture-matched runner and exercises PDF creation, saving, reopening, extraction, rendering, and immutable Pixmap storage. This covers all five abi3-py310 and all five CPython 3.14t artifacts; the sdist and PyEmscripten wheel have separate environment-specific gates.

Open, inspect, save

import pylopdf

doc = pylopdf.open("input.pdf")           # or pylopdf.open(stream=pdf_bytes)
print(doc.page_count)                     # len(doc) works too
print(doc.metadata["title"])
doc.set_metadata({"title": "Report", "author": "Alice"})

doc.save("out.pdf")
data = doc.tobytes()
doc.save("small.pdf", garbage=True, deflate=True, object_streams=True)
doc.save("locked.pdf", user_pw="secret", permissions=pylopdf.Permissions.PRINT)

Encrypted PDFs open with password= (or doc.authenticate() afterwards). pylopdf.peek_metadata(path) reads metadata and page count without fully parsing the document — useful when scanning large collections. Its optional max_file_size= rejects oversized path or byte input before parsing; the backward-compatible default is unbounded. Its repaired result and doc.is_repaired expose the bounded recovery of an incorrect final classic startxref; pylopdf also emits PylopdfWarning, and saving rewrites normalized cross-reference data. Pass limits=pylopdf.DocumentLimits.web() when processing untrusted files. It bounds file, structure, decompression, the rendering/extraction PDF snapshot, and interpreted text; inspect doc.complexity before heavy work and catch LimitError for controlled rejection. See Security for every budget.

page = doc[0]                             # 0-based; negative indices from the end
for page in doc:
    print(page.number, page.rect)

text = page.get_text()                    # plain text
words = page.get_text("words")            # (x0, y0, x1, y1, word, block, line, word_no)
layout = page.get_text("dict")            # blocks → lines → spans (bbox, size, font, flags)
hits = page.search_for("total")           # case-insensitive, list[Rect]

All coordinates are top-left origin display space — search results, layout, drawing and rendering all share the same coordinate system, including rotated pages.

Render

png = doc.render_page(0, dpi=300)                    # bytes (PNG)
pix = page.get_pixmap(scale=2)                       # RGBA8 pixels for NumPy / PIL
batch = doc.render_pages([0, 1, 2], scale=2, workers=4)
svg = doc.render_page_svg(0)

Edit

doc.delete_pages([1, 2])
doc.select([2, 0])                                   # keep/reorder (repeat = duplicate)
doc.new_page(); doc.copy_page(0, to=1)

merged = pylopdf.Document()
merged.insert_pdf(pylopdf.open("a.pdf"))
merged.insert_pdf(pylopdf.open("b.pdf"), from_page=0, to_page=2, start_at=0)

doc.set_toc([[1, "Chapter 1", 1], [2, "Section 1.1", 2]])
page.set_rotation(90)

Draw & annotate

page.insert_image((72, 72, 200, 200), filename="logo.png")   # JPEG passthrough / PNG alpha
page.insert_image(page.search_for("Approved")[0], stream=stamp_png)
page.insert_image((300, 72, 500, 200), pixmap=thumbnail, rotate=90)  # direct RGBA, clockwise rotation
page.show_pdf_page(page.rect, letterhead)                    # vector overlay; same-document src also works
page.insert_text((40, 40), "CONFIDENTIAL", fontsize=18, color=(1, 0, 0))
page.insert_text((40, 70), "社外秘", fontsize=18)             # auto-subset with pylopdf[cjk]
page.add_highlight_annot(page.search_for("important"))       # search & mark
page.add_link_annot(page.search_for("Example")[0], "https://example.com/")

Scanned PDFs, forms, Markdown

page.insert_ocr_text_layer(ocr_words)        # searchable PDFs from any OCR output
doc.set_form_field("customer", "Alice")      # native AcroForm appearance
doc.set_form_field("customer_ja", "山田 太郎")  # auto-subset with pylopdf[cjk]
md = doc.to_markdown()                       # RAG-ready Markdown

Continue with Ecosystem recipes for typesetting, PDF/A and digital signatures, or the migration guide if you come from pymupdf.