About extract links from PDF
A hyperlink in a PDF is not markup inside the text. It is a link annotation: a rectangle on a page with an action attached, and the text underneath it is entirely unrelated to where it goes. That separation is worth understanding, because it is why a PDF link can display one address and lead somewhere else, and why link text that looks like a URL is sometimes not a link at all. External links carry a URI action holding the target address; internal links carry a GoTo action naming a destination elsewhere in the document; some carry launch or remote-GoTo actions pointing at other files. This tool walks every page, collects every annotation of link type, resolves the action behind each one, and lists what it found alongside the text sitting under the rectangle — which is exactly the pairing you need for checking a long document before publication, or for auditing where a document you received wants to send you.
How to extract links from PDF
- 01
Load the PDF
Every page is scanned for link annotations and their associated actions.
- 02
Read the link table
Each entry gives the page, the target, the action type and the text under the link rectangle where text is present.
- 03
Check for mismatches
Compare displayed text against the actual target. A visible address that differs from the real destination is the classic phishing signature.
- 04
Export the list
Copy the results for a link check, a content audit or a migration inventory.
What this tool does
- External URLs from URI actions, deduplicated with a count of how many times each appears
- Internal destinations from GoTo actions, resolved to the target page number rather than left as a named destination
- Remote GoTo and launch actions surfaced separately, since they point outside the document
- The page each link appears on, so a broken link is easy to find and fix
- Text under the link rectangle captured where present, which is what makes a text-versus-target mismatch visible
- Named destinations resolved through the document names tree, including ones that point nowhere
- Link annotations with no action at all reported, since they are a common authoring mistake
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- This lists links; it does not test them. Whether a URL still resolves requires a network request, which would defeat the point of processing locally.
- Text that looks like a URL but was never made into a link annotation is not a link and is not listed. Use the text extractor if you want plain text addresses.
- A link rectangle can sit anywhere, so the text captured under it is the text it overlaps, which is usually but not always the intended label.
- JavaScript actions are reported by type rather than interpreted, since the target of a scripted navigation is not statically knowable.
- A document with no link annotations returns an empty list, which is normal for scanned and print-origin PDFs.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about extract links from PDF
Why do some visible web addresses not appear in the list?
Because they are only text. A PDF link is an annotation object layered over a region of the page, and text that reads like a URL has no link behaviour unless someone created that annotation — many authoring tools do it automatically, and many do not. If you need the addresses written in the copy, extract the text instead.
Can this find links that display one address and go somewhere else?
Yes, and that is one of the better reasons to run it. Because the annotation target is independent of the text beneath it, a PDF can show a familiar domain and lead anywhere. This report puts the captured text and the resolved target side by side so the mismatch is visible without clicking anything.
What is the difference between an external link and an internal destination?
An external link carries a URI action with an address and leaves the document. An internal link carries a GoTo action naming a destination — a page, often with a position and zoom — and moves within the document. Internal links are what a table of contents is built from, and this report resolves them to actual page numbers so you can see the ones pointing at pages that no longer exist.
Does this check whether the links still work?
No, by design. Testing a URL means requesting it, and nothing here touches the network — the whole document is parsed in your tab. Export the list and run it through a link checker if you need live status.
My table of contents links are broken. Will this show me why?
Often, yes. The two usual causes both appear here: a named destination that no longer resolves, which shows as an unresolved target, or a destination resolving to the wrong page after pages were inserted or removed. Seeing the resolved page numbers next to the entries usually makes the off-by-a-few pattern obvious.
Tools that pair with this one
- Extract PDF BookmarksExport a PDF outline tree, with nesting and target pages intact.
- Extract PDF AnnotationsTurn a marked-up PDF into a readable list of comments with author and page.
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- PDF InspectorA complete technical report on any PDF, produced without uploading it.
- Extract PDF MetadataExport both metadata layers a PDF carries — the info dictionary and the XMP packet.
- PDF Text ExtractorOne-pass extraction when you just need the words, fast.