Skip to content

Format · 8 min read · Updated 2026-09-19

What is a PDF?

A PDF records marks on a fixed canvas rather than the structure of a document. Understanding that one decision explains almost every PDF frustration.

Almost everyone uses PDFs daily and almost nobody is told what one is. It is not a word processor file with the editing removed, and it is not an image of a page. It is a program-like description of where ink goes, and once that clicks, the format stops seeming arbitrary.

What is a PDF?

01

A page description language, not a document model

Adobe created PDF in 1993 by taking PostScript — the page description language that drove professional typesetters — and removing its general-purpose programmability. PostScript was a real programming language: a printer had to execute an entire file from the start to discover what page 40 looked like. PDF kept the drawing model and threw away the loops and procedures, replacing them with a structure that can be opened at any page directly. In 2008 the specification was published as the open standard ISO 32000, which is why the viewer in your browser owes nothing to Adobe.

The word to hold onto is description. A PDF says: select this font at this size, move to this coordinate, show these glyphs, stroke a line from here to there, paint this image into this rectangle. It is an instruction list for making marks on a canvas of fixed dimensions. That is all it fundamentally is, and everything else in the format is scaffolding around it.

This is why PDF renders identically everywhere. Nothing is recomputed on arrival. A word processor file stores paragraphs and lets each machine work out where the lines break; a PDF stores the result of that calculation and lets no one revisit it. Fixed pagination is not a limitation bolted on, it is the deliverable.

02

What the file actually contains

Open a PDF in a text editor and the first bytes read %PDF- followed by a version number: the header. After it comes the body, a long sequence of numbered indirect objects. Each object is introduced by its number and generation, and holds a dictionary, an array, a number, a string or a stream. Pages are objects. Fonts are objects. The catalogue that names the page tree is an object. Objects refer to each other by number, so the body is a graph rather than a linear script.

Near the end sits the cross-reference table — in newer files a compressed cross-reference stream — which records the byte offset of every object in the file. Below it the trailer points at the catalogue and at that table. A viewer therefore reads the end of the file first, jumps straight to the objects it needs for page one, and never touches the rest. This is also why an interrupted download often produces a file that will not open at all: without an intact tail, nothing can be located.

The visible page lives in content streams, usually compressed. Inside are the drawing operators: text-positioning and text-showing operators, path construction and painting operators, graphics state changes for colour and line width, and operators to place external objects such as images. Alongside all this a PDF can carry annotations, form fields, bookmarks, embedded colour profiles, JavaScript, digital signatures and arbitrary file attachments. A PDF is much closer to a container than to a document.

03

The missing model, and everything it explains

Here is the fact that resolves most complaints about PDFs: the format has no paragraph, no heading and no table. Those concepts do not exist in the content stream. A table is some lines and some text that happen to sit in a grid. A heading is text that happens to be larger. A paragraph is a series of text-showing operations whose coordinates happen to descend the page at regular intervals. The relationship is visual coincidence, not recorded structure.

So editing a PDF is nothing like editing a document. Delete a word and nothing pulls the following words back, because nothing in the file knows they belong to the same sentence. Widen a margin and no line rebreaks. Editors that appear to reflow text are inferring a paragraph, rewriting the entire content stream and hoping their guess was right — which is why the result sometimes shifts in ways nobody asked for.

The same fact explains why PDF to Word is a reconstruction rather than a conversion. A converter receives positioned glyphs and must infer everything that matters: where paragraphs start, which lines form a table, what is a heading, whether two columns should be read side by side or one after the other. On a plain report it guesses well. On a designed brochure it cannot. And it explains why text extraction can return words in a surprising order — the extractor reads the content stream in the order the generator happened to write it, which need not be reading order.

There is an optional layer that supplies the missing structure: tagged PDF, a tree of marked-content roles saying this is a paragraph, this is a table cell, this is a heading. It is what screen readers need and what reliable reflow depends on. It is also entirely optional, and a great many PDFs in the world do not have it.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

Is a PDF an image of a page?

Usually not. A typical PDF exported from software contains real text as glyph references plus vector line work, which is why you can select a sentence and why it stays crisp at any zoom. Only scanned PDFs are genuinely images, and those contain nothing selectable until OCR adds a text layer.

Why can I not just edit a PDF like a Word document?

Because there are no paragraphs in the file to edit. The content stream records that a glyph sits at a coordinate, not that it belongs to a sentence, so removing a word leaves a gap rather than pulling the rest of the line back. Any editor offering reflow is reconstructing structure the format never stored.

Who controls the PDF format now?

It has been an open ISO standard since 2008, published as ISO 32000 and maintained through the ISO process rather than by Adobe alone. That is why independent readers, writers and browser engines can implement it fully without licensing anything.

Why does a corrupted PDF often fail to open at all?

Because the cross-reference table that locates every object sits at the end of the file and is read first. If the tail is missing or damaged, a viewer has no map of the object offsets, so even perfectly intact page content earlier in the file cannot be found without a repair pass that rebuilds the table by scanning.