Skip to content

Portable Document Format · .pdf

The PDF format

A fixed-layout page description format: the page you author is the page everyone sees. Open-standard since 2008, and the reason it is everywhere.

PDF describes a page rather than a document. Its job is to record exactly where every glyph, line and image sits on a fixed canvas, so that the file renders identically on a phone, a laptop and a commercial press. That single design decision explains both its dominance and every frustration people have with it.

The format itself

What PDF actually is

Adobe created PDF in 1993, deriving it from PostScript — the page description language that drove typesetters. It kept PostScript’s drawing model and dropped its general-purpose programmability, replacing it with a structure that can be opened at any page without executing the whole file first. In 2008 the specification was published as the open standard ISO 32000, which is why PDF readers and writers exist independently of Adobe.

Internally a PDF is a collection of numbered objects — dictionaries, arrays, streams — plus a cross-reference table that records the byte offset of each one. A viewer reads the trailer, jumps to the cross-reference table, and from there loads only the objects it needs. Page content itself lives in content streams: sequences of drawing operators that set a font, move to a coordinate, show text, fill a path, or paint an image.

Because the format carries both vector and raster content, embedded fonts, annotations, form fields, bookmarks and even arbitrary file attachments, a PDF is closer to a container than to a document file. What it deliberately does not carry is a reflowable text model. There is no notion of paragraphs flowing between pages, so editing a sentence does not push the rest of the document along — which is precisely why PDF editing feels awkward compared with a word processor.

Where it is the right answer

Common uses

  • 01

    Contracts, invoices and statements, where the recipient must see the exact layout that was issued

  • 02

    Forms that are filled in on screen, with real field values rather than typed-over images

  • 03

    Print-ready artwork handed to a commercial printer, with fonts embedded and colour defined

  • 04

    Long-term archiving of records under PDF/A, which forbids anything that might not render in decades

  • 05

    Scanned paper, where page images are wrapped in PDF and given a searchable text layer by OCR

  • 06

    Reports and documentation distributed for reading rather than further editing

Both sides

What PDF does well, and where it stops.

Strengths

  • Renders the same everywhere: pagination, fonts and positioning are fixed in the file, not recomputed by the reader
  • Mixes vector text and line work with raster images in one file, so quality scales to print resolution
  • Fonts can be embedded, removing any dependence on what the recipient has installed
  • Openly standardised, with mature independent implementations — including the ones running in your browser
  • Supports real structure on top of the page: bookmarks, links, tagged reading order, metadata and attachments

Limitations

  • Not an editing format. Revising text means rewriting content streams, and reflow simply does not exist in the model.
  • A scanned PDF contains only page images until OCR adds a text layer — the words are not searchable by default.
  • Accessibility is optional rather than inherent: an untagged PDF gives a screen reader little more than a visual jumble.
  • File size varies enormously for the same visible page, depending on whether images were downsampled and fonts subset.
  • Encryption and digital signatures are brittle under editing: any change to the bytes invalidates a signature by design.

Specification

Technical detail

The facts worth knowing before choosing PDF for a job, rather than after.

Full name
Portable Document Format
MIME type
application/pdf
Extension
.pdf
Introduced
1993, by Adobe
Standard
ISO 32000 (open standard since 2008)
Derived from
PostScript page description language
Layout model
Fixed page geometry — no reflow
Content types
Vector, raster, embedded fonts, annotations, form fields, attachments
Notable subsets
PDF/A (archival), PDF/UA (accessibility), PDF/X (print)

Questions about PDF

Why is editing a PDF so much harder than editing a Word file?

Because the format has no paragraph model. A PDF records that a glyph sits at a particular coordinate, not that it belongs to a sentence which belongs to a paragraph. Deleting a word therefore leaves a gap instead of pulling the following text back, and every editor has to reconstruct the missing structure by guesswork.

What is the difference between PDF/A, PDF/UA and PDF/X?

They are restricted profiles of the same format. PDF/A targets long-term archiving and forbids anything whose rendering depends on the outside world, such as unembedded fonts or external references. PDF/UA requires the tagging that assistive technology needs. PDF/X constrains colour and font handling for reliable commercial printing.

Why do two PDFs of the same document differ hugely in size?

Almost always images and fonts. A page scanned at high resolution and stored without downsampling can be many times the size of the same page exported from a word processor, and embedding a full font family rather than only the glyphs used adds further weight.

Can a PDF contain text that is not searchable?

Yes, and it is common. If the page is a photograph or scan of paper, the letters exist only as pixels. Running OCR adds an invisible text layer positioned over those pixels, which is what makes the document searchable without changing how it looks.