About PDF text diff
If you are used to diffing files, this is that tool pointed at PDFs. Granularity is fixed at word level, which is the right unit for prose: a line-based diff flags an entire paragraph because one article changed inside it, while a character-based diff shreds a renamed term into a confetti of single-letter edits. Word alignment marks exactly the tokens that were added or removed and leaves the rest as context. The presentation follows diff conventions rather than document-review conventions — additions and removals in a single aligned stream, unchanged text as context, so you can read the shape of the change quickly. The reason a PDF cannot simply be fed to a normal diff tool is that PDFs are binary and their bytes change wholesale on every save; the text has to be extracted and reconstructed into reading order first, which is what happens here before any alignment is attempted.
How to PDF text diff
- 01
Add the two files
The first is the "before" side of the diff, the second the "after".
- 02
Run the diff
Text is extracted from both, normalised, and aligned at word level.
- 03
Read the aligned output
Removals and additions are marked inline against unchanged context.
- 04
Export the report
Keep it as a record of what changed between two revisions.
What this tool does
- Word-level granularity fixed, avoiding both line-level over-reporting and character-level noise
- Diff-style aligned output with additions, removals and surrounding context
- Whitespace and hyphenation normalised so re-wrapping is not a difference
- Text reconstructed into reading order before alignment, which raw byte diffs cannot do
- Entirely local, so it works on source documents you cannot share
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- Word granularity means a single changed character shows as one word replaced by another; use the character setting on Compare PDF Files if you need to see which letter.
- No awareness of formatting, so styling and layout differences are not reported.
- Scanned documents produce no text and therefore no diff without OCR first.
- Moved blocks may be reported as a deletion in one place and an insertion in another rather than as a move.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about PDF text diff
Why can't I just run diff on the two PDF files?
Because PDFs are binary containers whose bytes change on every save, including object ordering and timestamps. You would see the whole file as different. The text has to be extracted first.
Why is word the right granularity?
It matches how prose changes. Line diffs mark a whole paragraph when one word moved; character diffs turn a renamed term into dozens of tiny edits. Word level shows the edit as the edit.
Are moved paragraphs detected as moves?
Not as such. The alignment reports them as a removal and an addition. Reading the two together usually makes the move obvious, but it is not labelled as one.
How is this different from Compare PDF Text?
The engine and mode are the same; the output framing differs. This page gives you diff-style aligned output, while Compare PDF Text is organised for reviewing and approving a revision.
Tools that pair with this one
- Compare PDF FilesTwo versions in, a clear list of differences out.
- Compare PDF TextOnly the words, so a reformatted document still reads as unchanged.
- Visual PDF ComparisonCompare what the page looks like, not just what it says.
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- PDF Text CleanerExtraction plus the clean-up pass you would otherwise do by hand.
- PDF InspectorA complete technical report on any PDF, produced without uploading it.