Portable Document Format · .pdf
The PDF format
A fixed-layout page description format: the page you author is the page everyone sees. Open-standard since 2008, and the reason it is everywhere.
PDF describes a page rather than a document. Its job is to record exactly where every glyph, line and image sits on a fixed canvas, so that the file renders identically on a phone, a laptop and a commercial press. That single design decision explains both its dominance and every frustration people have with it.
The format itself
What PDF actually is
Adobe created PDF in 1993, deriving it from PostScript — the page description language that drove typesetters. It kept PostScript’s drawing model and dropped its general-purpose programmability, replacing it with a structure that can be opened at any page without executing the whole file first. In 2008 the specification was published as the open standard ISO 32000, which is why PDF readers and writers exist independently of Adobe.
Internally a PDF is a collection of numbered objects — dictionaries, arrays, streams — plus a cross-reference table that records the byte offset of each one. A viewer reads the trailer, jumps to the cross-reference table, and from there loads only the objects it needs. Page content itself lives in content streams: sequences of drawing operators that set a font, move to a coordinate, show text, fill a path, or paint an image.
Because the format carries both vector and raster content, embedded fonts, annotations, form fields, bookmarks and even arbitrary file attachments, a PDF is closer to a container than to a document file. What it deliberately does not carry is a reflowable text model. There is no notion of paragraphs flowing between pages, so editing a sentence does not push the rest of the document along — which is precisely why PDF editing feels awkward compared with a word processor.
Where it is the right answer
Common uses
- 01
Contracts, invoices and statements, where the recipient must see the exact layout that was issued
- 02
Forms that are filled in on screen, with real field values rather than typed-over images
- 03
Print-ready artwork handed to a commercial printer, with fonts embedded and colour defined
- 04
Long-term archiving of records under PDF/A, which forbids anything that might not render in decades
- 05
Scanned paper, where page images are wrapped in PDF and given a searchable text layer by OCR
- 06
Reports and documentation distributed for reading rather than further editing
Both sides
What PDF does well, and where it stops.
Strengths
- Renders the same everywhere: pagination, fonts and positioning are fixed in the file, not recomputed by the reader
- Mixes vector text and line work with raster images in one file, so quality scales to print resolution
- Fonts can be embedded, removing any dependence on what the recipient has installed
- Openly standardised, with mature independent implementations — including the ones running in your browser
- Supports real structure on top of the page: bookmarks, links, tagged reading order, metadata and attachments
Limitations
- Not an editing format. Revising text means rewriting content streams, and reflow simply does not exist in the model.
- A scanned PDF contains only page images until OCR adds a text layer — the words are not searchable by default.
- Accessibility is optional rather than inherent: an untagged PDF gives a screen reader little more than a visual jumble.
- File size varies enormously for the same visible page, depending on whether images were downsampled and fonts subset.
- Encryption and digital signatures are brittle under editing: any change to the bytes invalidates a signature by design.
Specification
Technical detail
The facts worth knowing before choosing PDF for a job, rather than after.
- Full name
- Portable Document Format
- MIME type
- application/pdf
- Extension
- Introduced
- 1993, by Adobe
- Standard
- ISO 32000 (open standard since 2008)
- Derived from
- PostScript page description language
- Layout model
- Fixed page geometry — no reflow
- Content types
- Vector, raster, embedded fonts, annotations, form fields, attachments
- Notable subsets
- PDF/A (archival), PDF/UA (accessibility), PDF/X (print)
Tools that apply
Converting PDF in your browser
Each listed MyPDFilles conversion runs on your device: the file is read from your disk, converted in the tab and handed back as a download. Educational references to external software are not covered by that claim.
- PDF to JPG
PDF to JPG
Render each page of a PDF as a JPG image, at the resolution you pick.
- PDF to PNG
PDF to PNG
Render PDF pages as white-background PNG images with lossless pixel encoding.
- PDF to WebP
PDF to WebP
Export PDF pages as WebP, then compare size, appearance and compatibility.
- PDF to text
PDF to Text
Recover the words from a PDF — read out of its text layer, not guessed.
- PDF to Word
PDF to Word
An editable Word document rebuilt from your PDF’s text and positions.
- Compress PDF
Compress PDF
Make a PDF smaller by re-encoding its images and cleaning up its internals.
- Merge PDF
Merge PDF
Combine several PDFs into a single document, in the order you choose.
- Split PDF
Split PDF
Cut one PDF into several documents — by range, by interval or page by page.
Questions about PDF
Why is editing a PDF so much harder than editing a Word file?
Because the format has no paragraph model. A PDF records that a glyph sits at a particular coordinate, not that it belongs to a sentence which belongs to a paragraph. Deleting a word therefore leaves a gap instead of pulling the following text back, and every editor has to reconstruct the missing structure by guesswork.
What is the difference between PDF/A, PDF/UA and PDF/X?
They are restricted profiles of the same format. PDF/A targets long-term archiving and forbids anything whose rendering depends on the outside world, such as unembedded fonts or external references. PDF/UA requires the tagging that assistive technology needs. PDF/X constrains colour and font handling for reliable commercial printing.
Why do two PDFs of the same document differ hugely in size?
Almost always images and fonts. A page scanned at high resolution and stored without downsampling can be many times the size of the same page exported from a word processor, and embedding a full font family rather than only the glyphs used adds further weight.
Can a PDF contain text that is not searchable?
Yes, and it is common. If the page is a photograph or scan of paper, the letters exist only as pixels. Running OCR adds an invisible text layer positioned over those pixels, which is what makes the document searchable without changing how it looks.
Compare with
Related formats
DOCX
.docxA ZIP container of XML describing a reflowable document. Pagination is not stored — it is calculated when the file is opened.
JPG
.jpgThe format photographs live in. Lossy block-based compression trades detail you are unlikely to notice for files small enough to send.
PNG
.pngLossless, with a genuine alpha channel. The correct home for screenshots, logos, UI and anything with hard edges you cannot afford to soften.
TIFF
.tifNot one encoding but a container that can hold many. The long-standing standard in scanning, fax and prepress, and unreadable in a browser.