About extract PDF annotations
When a document comes back from review, the feedback is not in the document text — it is in a parallel layer of annotation objects, one per page, each with its own subtype, rectangle, author, timestamp and contents. That layer is genuinely useful data trapped in an awkward container: to read it in a viewer you click through a comments panel, and to act on it you retype it into a task list. This tool reads the annotation objects directly and produces the list, ordered by page, with the author and date each comment carries and the quoted text a highlight or strikeout covers. It handles the whole family of markup subtypes — text notes, highlights, underlines, strikeouts, squiggles, free text, ink drawings, stamps, shapes — and it distinguishes them, because "seven comments" and "seven highlights and no comments" mean rather different things about a review.
How to extract PDF annotations
- 01
Load the marked-up PDF
Every page annotation is read locally, including reply threads where a reviewer answered another reviewer.
- 02
Read the comment list
Comments are grouped by page with subtype, author, date and contents, plus the text a markup annotation covers.
- 03
Filter by what matters
Sort or scan by author when you need one reviewer's pass, or by subtype when you want only the substantive notes.
- 04
Export the record
Copy the list into a ticket tracker, a change log or a reply to the reviewer.
What this tool does
- Every markup subtype captured: Text notes, Highlight, Underline, StrikeOut, Squiggly, FreeText, Ink, Square, Circle, Line, Polygon, Stamp and Caret
- Author from the /T entry and timestamp from /M, reported per annotation where the producer recorded them
- Comment body text extracted from /Contents, including multi-line notes
- Quoted page text under highlight, underline and strikeout annotations, so a note has its context
- Reply threads kept with their parent annotation rather than flattened into the list
- Page number and position for every annotation, so each one can be found again in the original
- Counts by subtype and by author, which is the quickest summary of what a review actually contained
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- Many PDFs carry no annotations. An empty result means the document has not been marked up, not that extraction failed.
- Author and date are only as good as the producer wrote them. Some tools omit the author entirely, and this report leaves the field blank rather than guessing.
- Text captured under a highlight is derived from the page text within the annotation rectangle, so an imprecisely drawn highlight yields imprecise quoted text.
- Ink drawings, stamps and shapes are visual. They are listed with their type, page and any text label, but a drawn circle has no textual content to report.
- Annotations flattened into the page — burned into the content stream rather than left as objects — cannot be recovered, because they are no longer annotations.
- Redaction annotations are reported as annotations; this tool neither applies nor reveals redacted content.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about extract PDF annotations
Why does my marked-up PDF show no annotations?
The most common reason is flattening. Some export and print-to-PDF paths draw the markup permanently into the page content and discard the annotation objects, which makes the comments visible but no longer readable as data. If that has happened, the original annotated file is the only source. The other possibility is simply that the markup you are looking at is page content authored deliberately rather than review annotations.
Can I see who made each comment and when?
Where the producer recorded it, yes — each annotation can carry an author string and a modification date, and both are reported per comment. Tools vary in how diligently they set these, and a reviewer working in an application that does not ask for a name will leave the author blank.
Does this include the text that was highlighted?
Yes. A highlight annotation does not itself store the text it covers, so the quoted text is derived by reading the page text within the annotation rectangle. That works well for ordinary prose and less well where a highlight was drawn loosely across a column boundary or a table.
How is this different from extracting the document text?
Completely different layers. Text extraction returns what the document says. This returns what reviewers said about it — a separate set of objects that sit on top of the pages and that no text extractor will ever return.
Are reply threads preserved?
Yes. A reply is an annotation with an /IRT entry pointing at the annotation it answers, so the thread structure exists in the file and is kept in the output rather than being flattened into a sequence of unconnected comments.
Tools that pair with this one
- Extract Links from PDFPull every URL and internal destination out of a PDF, with the page each one is on.
- Extract PDF BookmarksExport a PDF outline tree, with nesting and target pages intact.
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- PDF InspectorA complete technical report on any PDF, produced without uploading it.
- Compare PDF FilesTwo versions in, a clear list of differences out.
- Flatten PDF FormBake the current answers into the page and drop the fillable layer.