Skip to content

Developer Tools

PDF Structure Inspector

See how the document is wired together, from trailer to page tree.

Processed locally in your browser. Your file is not uploaded to any server. How this works

  1. 01Add your file
  2. 02PDF Structure Inspector
  3. 03Download

Runs in your browser · nothing uploaded

About PDF structure inspector

Reading a PDF starts at the end of the file. A reader seeks to the last line, finds the byte offset of the cross-reference index, reads that index to learn where every object lives, then follows the trailer to the document catalogue and from there down the page tree to the pages themselves. That chain is the file structure, and when a PDF misbehaves in ways that have nothing to do with how it looks — a viewer that refuses to open it, a web server that cannot stream it page by page, an editor that shows an older version of a document than the one you saved — the answer is almost always somewhere in that chain. This tool walks it and shows you what it found: the catalogue keys, the shape of the page tree, whether the cross-reference index is a classic table or a compressed stream, whether the file is linearised for progressive loading, and how many times it has been appended to by incremental update.

How to PDF structure inspector

  1. 01

    Load the file

    Drop the PDF in. Parsing starts from the trailer, exactly as a reader would.

  2. 02

    Read the chain

    Trailer, cross-reference index, catalogue, page tree — each step is reported with the keys actually present.

  3. 03

    Check the structural flags

    Xref type, linearisation, object streams and update generations are called out separately because each has practical consequences.

What this tool does

  • Trailer keys and the resolved document catalogue, with the entries that are present rather than a fixed template
  • Page tree shape: node depth, kids-per-node and whether the tree is balanced or a single flat node
  • Cross-reference index type: classic table, compressed xref stream, or a hybrid file carrying both
  • Linearisation detection, including whether the hint tables are actually present
  • Incremental update generations, counted from the chain of /Prev pointers
  • Object stream usage, which is what makes an object count higher than the raw xref entries suggest

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • Structural detail is descriptive. A file can be structurally unusual and still render perfectly in every viewer you care about.
  • Repairable damage is reported as damage: if the xref offsets are wrong, the report says the index was rebuilt by scanning rather than pretending it was valid.
  • Encrypted documents require the password before the object graph can be resolved.
  • The number of incremental updates is a lower bound. A producer that rewrites the file completely leaves no trace of earlier revisions.

How your file is handled

This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about PDF structure inspector

What is the difference between an xref table and an xref stream?

Both index the objects in the file. A classic cross-reference table is plain text — a list of twenty-byte entries giving each object byte offset — and is what PDF 1.0 to 1.4 use. An xref stream, introduced with PDF 1.5, stores that same index as a compressed binary stream, which is smaller and allows objects themselves to be packed into object streams. Very old readers understand only the table form, which is why some producers write hybrid files containing both.

What does linearization actually do?

A linearised PDF, marketed as Fast Web View, is reorganised so that everything needed to draw the first page comes first in the byte stream, followed by hint tables telling a reader where the remaining pages are. Over a connection that supports byte ranges, a viewer can show page one before the file has finished downloading, and can jump to page fifty by fetching only that region. The file renders identically either way; linearisation is purely about delivery.

My file shows several incremental update generations. Is that bad?

Not inherently — it is how PDF supports appending changes without rewriting the whole file, and it is required for signing, because an incremental save leaves earlier signed bytes untouched. The practical consequences are that the file carries its own history, so previous versions of edited content may still be present in the bytes, and the file is larger than a clean rewrite would be.

Why does the object count differ from the number of xref entries?

Because objects can be packed inside object streams. The xref stream points at the container, and the container holds many compressed objects. The total count follows the objects; the entry count follows the containers.

Should I use this or the full PDF Inspector?

Use the full inspector when you want a survey of content — fonts, images, colour, metadata. Use this one when the question is about the file as a file: why a reader rejects it, why it will not stream, or how many times it has been appended to.

Read more about this