Skip to content

Security · 7 min read · Updated 2026-09-19

What is PDF metadata?

Every PDF carries fields describing itself. Most were filled in by software you never configured, and some of them say more than you intended.

Metadata is data about the document rather than content of the document. It never appears on the page, which is exactly why it goes unexamined. Two separate mechanisms store it in a PDF, and a file will frequently contain both — sometimes disagreeing with each other.

What is PDF metadata?

01

Two places metadata lives

The older mechanism is the document information dictionary, referenced from the trailer. It holds a small, fixed set of named entries: /Title, /Author, /Subject, /Keywords, /Creator, /Producer, /CreationDate and /ModDate. This is what most viewers show in a document properties panel, and it is the layer that tools have read and written since the early days of the format.

The newer mechanism is XMP: Extensible Metadata Platform packets, embedded in the PDF as streams containing RDF/XML. Because it is XML, XMP is open-ended — it can carry the same basic fields plus whole vocabularies for rights, provenance, colour management, document identifiers, and application-specific data that a particular tool decided to record. PDF/A relies on XMP to declare its conformance claim, and many production workflows depend on it.

Both can be present simultaneously, and nothing forces them to agree. Editing the information dictionary in one tool while an XMP packet elsewhere retains the old author is a common and easily missed inconsistency. It is also why removing metadata properly means addressing both layers: clearing the visible properties panel while leaving an XMP packet untouched achieves very little.

02

What the fields actually mean

Two entries are routinely misread. /Creator names the application that created the original content — the word processor, layout program or CAD tool you authored in. /Producer names the software that wrote the PDF itself: the export library, print driver or conversion utility. A document written in one program and exported through another therefore records both, and the pair is unusually informative because it describes your toolchain rather than just one program.

Worse, these strings are frequently verbose. Producers commonly write their exact product name and version number. That is a precise statement of what software, at what release, was installed on the machine that made the file — the sort of detail that is mildly interesting in a support ticket and less welcome in a document sent to a counterparty.

/Author defaults to whatever account name or registered user the authoring software knows, which is often a full real name nobody consciously entered. /CreationDate and /ModDate are timestamps, typically including a UTC offset, so they can reveal when a document was actually prepared — including that a document dated Friday was written at two in the morning on Sunday, or that a contract described as freshly drafted was created months earlier. /Title is frequently left as the original filename or as a stray heading, which is how internal working titles end up travelling with public files.

03

The genuine leak risks

Names are the most common exposure. An author field carrying a real name, or a series of names accumulated through revisions, tells a recipient who worked on a document even when the document itself is presented as institutional. In some XMP vocabularies a list of contributors or prior authors persists across edits.

Local file paths are the sharpest. Some producers and some embedded content record the full path of the source file or of linked resources, which can disclose usernames, internal drive letters, server names, project codenames and organisational structure. A path such as a departmental share under a named user reveals more about an organisation than most documents intend to.

Software and version disclosure has a security dimension as well as a privacy one: publishing the exact release of the tool that produced a document tells anyone interested which known issues that release has. Revision traces are another category — some workflows leave document identifiers and version history in XMP, so successive drafts of the same file can be linked together even when the visible content was rewritten.

None of this is visible on the page, and none of it is removed by printing to a new PDF in every case — some producers faithfully carry XMP forward. Inspecting what a file actually contains before sending it is the only reliable approach, and stripping metadata is the fix once you know what is there. The trade-off is worth stating: metadata also serves legitimate purposes, since archival identification, accessibility declarations and print workflows all depend on it, so blanket removal is not always the right answer for files headed into a managed pipeline.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

What is the difference between /Creator and /Producer?

/Creator records the application the content was authored in, such as a word processor or layout tool. /Producer records the software that actually wrote the PDF bytes — an export library, print driver or converter. Together they describe your toolchain, often including exact version numbers.

Does removing metadata change the visible document?

No. Metadata never appears on the page, so stripping it leaves the rendered content identical. The caveats are functional rather than visual: a PDF/A conformance claim lives in XMP and would be lost, and managed print or archival workflows may expect fields to be present.

Can a PDF reveal file paths from my computer?

Yes, it happens. Some producers and some embedded or linked resources record full source paths, which can expose usernames, drive letters, server names and internal project names. Inspecting the metadata before sharing a file is the only way to know whether yours does.

Why does my PDF show a different author after I edited it?

Usually because the two metadata layers disagree. A tool may have updated the document information dictionary while an XMP packet still holds the original values, or vice versa, so different viewers read different sources and report different authors.