Skip to content

Format · 7 min read · Updated 2026-09-19

What is PDF/A?

An archival subset of PDF that forbids anything whose appearance depends on the outside world. Same format, deliberately fewer freedoms.

PDF/A is not a different file type. It is ordinary PDF with rules attached, designed so that a document opened in forty years still renders as it did on the day it was filed. Every rule it adds exists to remove a dependency on something outside the file.

What is PDF/A?

01

Self-containment as the single design principle

An ordinary PDF is allowed to lean on its environment. It may reference a font by name and assume the reader has it. It may describe colour in a way that depends on the display. It may link to a file on a network share. Each of those is a bet that something outside the file will still be there and still behave the same way. Archiving is the business of not making bets.

PDF/A therefore requires that fonts be embedded — every glyph actually used must travel inside the file, so nothing is substituted decades later when the original typeface is unavailable. It requires device-independent colour, meaning colour must be specified in a way that is unambiguous rather than left to whatever the rendering device assumes, typically by embedding an ICC profile. It forbids external dependencies: no references to files outside the document, no reliance on fetched resources.

It also forbids encryption outright. This surprises people, but it follows directly: a file whose bytes can only be read with a key is a file that becomes unreadable the moment the key is lost, and an archive cannot accept that. For the same reason PDF/A bars JavaScript, launch actions and embedded multimedia that would need an external player — anything whose behaviour is not fully determined by the file itself.

Finally it requires XMP identification metadata: an embedded XMP packet declaring which PDF/A part and conformance level the file claims. That declaration is what makes a file self-describing and what a validator checks first.

02

The parts: A-1, A-2 and A-3

PDF/A-1 arrived in 2005 and is based on PDF 1.4. Being the oldest and most restrictive, it is also the most conservative choice: it predates and therefore excludes later PDF features such as transparency and JPEG 2000 compression. Files that comply with it are constrained but very unlikely to trouble any implementation.

PDF/A-2 followed in 2011, built on the later PDF 1.7 feature set. It permits things A-1 forbade — transparency, layers, JPEG 2000 image compression, and embedded OpenType fonts — because those features had become universally implemented and no longer represented a rendering risk. For most new archiving work A-2 is the practical default.

PDF/A-3 is A-2 plus one significant change: it allows arbitrary files to be embedded as attachments. That means a PDF/A-3 file can carry its own source alongside the rendered page — a spreadsheet next to the report generated from it, or a machine-readable invoice next to its human-readable form, which is how electronic invoicing formats use it. The relaxation is genuinely useful and also genuinely a loosening: the attached file itself is not required to be archival, so the guarantee covers the PDF pages, not necessarily what is stapled to them.

03

Conformance levels: b, a and u

Each part offers conformance levels, and they are about how much of the document meaning is preserved, not about how strictly the rules were followed. Level b — basic — guarantees reliable visual reproduction. The page will look right. It says nothing about whether the text underneath is intelligible to anything other than an eye.

Level a — accessible — is substantially more demanding. It requires the document to be tagged with logical structure, so the reading order, headings, lists and table relationships are recorded, and it requires that text map to Unicode. This is what a screen reader needs, and it is why level a is difficult to achieve retroactively: the structure has to come from the authoring process, because it cannot be reliably inferred from a finished page.

Level u sits between them, introduced with PDF/A-2. It requires Unicode mapping for all text but not full tagging. In practice that means the text can be searched, extracted and copied correctly — you get meaning out of the file rather than just pixels — without the full accessibility apparatus. A file described as PDF/A-2b is therefore visually archival; PDF/A-2u is visually archival and textually trustworthy; PDF/A-2a is both of those plus structurally accessible.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

Why does PDF/A forbid encryption?

Because an archive must remain readable without any external secret. Encrypted bytes are unreadable without the key, and over the decades an archive is meant to span, keys get lost, rotated or forgotten. A file that depends on one is not self-contained, which is the entire premise of the standard.

Does converting a file to PDF/A change how it looks?

It can. If fonts were not embedded, conversion has to embed them or substitute something available, and a substituted typeface with different metrics shifts line widths. Transparency also has to be flattened for PDF/A-1. Always compare the converted file against the original rather than assuming the appearance survived.

Which PDF/A version should I use?

PDF/A-2 suits most new archiving because it allows transparency and modern compression while keeping the archival guarantees. Choose A-1 when a receiving institution specifically requires it, and A-3 only when you genuinely need to embed source files alongside the pages.

Is a PDF/A file still a normal PDF?

Yes. Any PDF reader opens it without knowing anything about the standard, because PDF/A is a restricted profile rather than a separate format. The conformance claim lives in an XMP metadata packet that validators read and ordinary viewers simply ignore.