About extract PDF metadata
PDFs store metadata twice, and the two copies regularly disagree. The document information dictionary is the older mechanism: a simple set of keys in the trailer holding title, author, subject, keywords, creator, producer and the creation and modification dates. The XMP packet is the newer one: an RDF/XML document embedded as a metadata stream, carrying the same properties in standardised namespaces plus whatever else a producer wants to record — rights, provenance, a document identifier, custom application fields. Editors are inconsistent about updating both, so it is entirely routine for a file to show one title in a viewer's properties panel and a different one in its XMP. This tool exports both layers in full, side by side, for a data-oriented purpose: building an inventory across a document set, auditing what a file discloses before it leaves the building, or capturing provenance for a records system.
How to extract PDF metadata
- 01
Load the PDF
The information dictionary is read from the trailer and the XMP stream from the catalogue.
- 02
Compare the two layers
Both are exported in full, with fields that disagree between them called out.
- 03
Export the record
Take it as JSON for an inventory or a pipeline, or as a flat key-value list for a spreadsheet.
- 04
Act on what you find
Use the metadata editor to correct fields, or the metadata remover to strip them before publication.
What this tool does
- Complete document information dictionary, including non-standard keys a producer added
- Full XMP packet exported, with namespace prefixes preserved rather than flattened away
- Creation and modification dates parsed from PDF date syntax, with time zone offsets kept
- Producer and creator reported separately, which is what reveals the actual toolchain a document passed through
- Custom and application-specific properties retained instead of filtered to a known list
- Document instance and original document identifiers, which is how revisions of one document are traced
- Disagreements between the info dictionary and XMP flagged explicitly
- Export as JSON or as a flat key-value list
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- This exports metadata; it does not edit or remove it. Use the metadata editor or remover for that.
- Metadata is unverified, self-declared text. An author field says what someone typed, which may be wrong, stale or deliberately misleading.
- Absent fields are absent. Many files carry only a producer string, and an empty title is common rather than anomalous.
- Page-level and object-level metadata streams, which a few producers write, are not merged into the document-level export.
- XMP is exported as recorded. A malformed packet is reported as malformed rather than silently repaired.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about extract PDF metadata
How is this different from the metadata viewer?
The viewer is for reading — it presents the fields in a friendly panel for a person answering a question about one document. This is for exporting: both layers in full, including custom and non-standard keys, in JSON or key-value form, for inventories, audits and pipelines across many files.
Why do the info dictionary and XMP disagree in my file?
Because they are separate stores that a producer has to keep in sync deliberately, and many do not. An editor that updates the title in one place and not the other leaves exactly this mismatch, which is why a viewer's properties panel and a search index built from XMP can genuinely show different titles for the same file.
What does this reveal that I might not want published?
More than people expect. Author names and usernames, the software and version that produced the file, creation and revision timestamps that date a document more precisely than its content does, and occasionally original filenames or paths. If the document is going outside your organisation, look at this first and strip what you would not want quoted.
Can I trust the creation date?
Only as far as you trust its source. Dates are written by the producing software from the local clock and can be edited afterwards with no trace. They are useful evidence and not proof, which matters whenever a date is doing real work in a dispute.
What is the document ID for?
XMP records an original document identifier and an instance identifier. The original stays constant across revisions of the same document while the instance changes on each save, so together they let a system recognise two files as versions of one work rather than as unrelated documents.
Tools that pair with this one
- PDF Metadata ViewerRead the fields a PDF carries about itself, including the raw XMP.
- PDF Metadata RemoverClear selected document properties before sharing; this is not secure anonymisation.
- Edit PDF MetadataSet a document’s properties deliberately instead of inheriting them.
- PDF InspectorA complete technical report on any PDF, produced without uploading it.
- Extract PDF AttachmentsRecover the files hidden inside a PDF as embedded attachments.
- PDF/A CheckerCheck the archival indicators in a PDF — clearly labelled as indicators, not certification.