.text is flattened; .pages is not

result.text (Lesson 2) is every page's text joined into one string, in reading order, with page boundaries thrown away. That's fine when you don't care which page a sentence came from. The moment you do, page numbers for citations, per-page chunking for a vector index, skipping a known-bad page, you need result.pages instead.

ParsedPage, field by field

Each entry in result.pages is a ParsedPage:

FieldWhat it is
page_num1-indexed page number
width, heightPage size in points (72 points = 1 inch; 595x842 is A4)
textThis page's own text, the slice result.text is built from
text_itemsA list of TextItem, one per run of text, each with its own bounding box and font
markdownPopulated only with output_format="markdown" (Lesson 3)

text_items is the layer underneath .text: instead of one flat string, you get where each piece of text sits on the page (x, y, width, height, in the same point units as the page itself) and what font/size it was set in. This is the foundation later lessons build on: form fields and annotations (Lesson 6, 9) come with their own rects in this same coordinate space.

The code, piece by piece

for page in result.pages:
print(page.page_num, page.width, page.height, len(page.text))

Straightforward iteration. Note page.text here is that page's own text, not the whole document's, this document only has one page so they happen to match, but on a multi-page PDF they wouldn't.

first_item = page.text_items[0]
print(first_item.text, first_item.x, first_item.y, first_item.font_name, first_item.font_size)

The document's title, as a single TextItem, positioned near the top of the page (y=56.7, close to 0 which is the page's top edge in LiteParse's top-left-origin coordinate system) in a 10pt monospaced font.

page_one = result.get_page(1)
missing = result.get_page(99)

get_page(n) is a convenience lookup by 1-indexed page number, an alternative to result.pages[n - 1]. It returns None instead of raising when the page doesn't exist, worth checking for if page numbers come from somewhere outside your control (like user input).

Checkpoint

  • result.pages: a list of ParsedPage, one per page, in order.
  • page.text: that page's own text, distinct from result.text (the whole document, flattened).
  • page.text_items: the finer-grained layer, one TextItem per text run, with its own bounding box and font info.
  • result.get_page(n): a 1-indexed, None-safe page lookup.

If anything here still feels unclear, ask before moving to Lesson 5.