Skip to content

PDF OCR

Make PDF Searchable

Keep the scan exactly as it looks; make Ctrl+F start working.

Processed locally in your browser. Downloads a recognition engine once; your file is still never uploaded. How this works

  1. 01Add your file
  2. 02Make PDF Searchable
  3. 03Download

Runs in your browser · downloads an engine file once

About make PDF searchable

This page is for one specific problem: an archive of scans that nobody can search. The fix is a text layer, and the important thing to understand is that it is additive. Your scanned image stays exactly where it is, untouched, as the visible content of the page. Behind it, invisible, we place the recognised words at the coordinates where those words appear in the image — rendered in a transparent text mode so they draw nothing. Visually the output is indistinguishable from the input, which is what makes this safe for records where the appearance of the page is the record. Functionally the file changes character completely: Ctrl+F finds terms, a text selection highlights the right area, copy yields words, and desktop search indexers can finally see inside the document. Nothing is retyped and nothing is re-laid-out, so there is no risk of an OCR error silently changing what the page shows.

How to make PDF searchable

  1. 01

    Add the scan

    Drop in the scanned PDF whose contents you cannot currently search.

  2. 02

    Pick the language

    Choose from the seven trained models so the recogniser knows which letter shapes and word patterns to expect.

  3. 03

    Run recognition

    Each page is rendered, recognised on your CPU, and given its hidden layer. Long files take minutes.

  4. 04

    Save and test

    Download the result and press Ctrl+F. The page looks the same; the search now works.

What this tool does

  • Original scanned images preserved byte-for-byte as the visible page
  • Invisible text layer positioned word by word behind the image
  • Output is searchable by readers and by desktop search indexers
  • Recognition errors never alter what the page displays, only what it matches
  • Runs on your own machine; the scan is never transmitted

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • Only English, French, Spanish, German, Italian, Portuguese and Arabic have models here.
  • The hidden layer inherits any recognition errors, so a misread word will not be found by search even though the page reads correctly.
  • File size grows, because the text layer is added on top of the images that were already there.
  • Handwritten pages will not gain a useful text layer; the models are trained on print.

How your file is handled

This tool runs inside this browser tab, but it first downloads a recognition engine and language model — static files, fetched once and then cached by your browser. Your document is never part of that request: the engine comes down to your device, and your file stays on it. You can verify this in the Network panel, where you will see the engine assets download and no upload of your document.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about make PDF searchable

Will the page look any different afterwards?

No. The scanned image remains the visible content. The added text is drawn in an invisible rendering mode, so it occupies the right coordinates while displaying nothing.

What if OCR misreads a word?

The visible page is unaffected — you still see the original scan. The consequence is narrower: that particular word will not be found by a search, since the hidden layer holds the misreading.

Why is the output file bigger than the input?

Because nothing was removed. You now have the original images plus a text layer. If size matters, run the result through Compress PDF.

How is this different from OCR PDF?

OCR PDF is the general tool with a text-only option. This page locks the output to searchable-PDF, because that is the only outcome that serves an archive where the scan must keep its exact appearance.

Read more about this