Skip to content

PDF OCR

PDF to Searchable Text

When you want the words out of a scan, not a prettier scan.

Processed locally in your browser. Downloads a recognition engine once; your file is still never uploaded. How this works

  1. 01Add your file
  2. 02PDF to Searchable Text
  3. 03Download

Runs in your browser · downloads an engine file once

About PDF to searchable text

There are two useful outcomes for OCR on a scan, and this is the other one. Instead of putting a hidden layer behind the images and keeping the PDF, this discards the pictures entirely and returns the recognised words as plain text. That is the right choice when the text is the deliverable and the page appearance is not: loading scanned records into a database, indexing a document set for search in your own application, feeding historical material to a language model, or simply editing content that only exists on paper. Practically it is also far lighter — a 40 MB scanned report becomes a text file measured in kilobytes, which is the difference between something you can grep and something you cannot. The trade is explicit: you lose the visual record. Keep the original scan if the page itself has evidentiary value, and use Make PDF Searchable instead when the appearance must be preserved.

How to PDF to searchable text

  1. 01

    Add the scanned PDF

    Drop in the file whose text you need extracted rather than hidden.

  2. 02

    Choose the language

    Select from the seven trained models to match the document.

  3. 03

    Run recognition

    Each page is rendered and recognised on your CPU; expect a few seconds per page.

  4. 04

    Save the text

    Copy it or download a .txt. Page boundaries are marked so you can still cite by page.

What this tool does

  • Plain-text output, typically thousands of times smaller than the scanned original
  • Page boundaries marked so citations by page number remain possible
  • Result is directly greppable, indexable and editable
  • Recognition on your own machine, suitable for confidential archives
  • Seven trained languages, cached after the first download

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • The visual record is discarded — keep the original PDF if the appearance of the page matters.
  • Recognition errors land directly in the text you will use, so proofread anything consequential.
  • Multi-column scans and tables can come out in the wrong reading order.
  • Only English, French, Spanish, German, Italian, Portuguese and Arabic are available, and handwriting is not reliably read.

How your file is handled

This tool runs inside this browser tab, but it first downloads a recognition engine and language model — static files, fetched once and then cached by your browser. Your document is never part of that request: the engine comes down to your device, and your file stays on it. You can verify this in the Network panel, where you will see the engine assets download and no upload of your document.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about PDF to searchable text

When should I choose this over Make PDF Searchable?

Choose this when the words are the product — indexing, editing, data loading. Choose the searchable PDF when the page must keep looking exactly as it does and search is an addition rather than the goal.

Why is the output so much smaller?

Because the images are gone. Scanned page images are nearly all of a scanned PDF's size; the recognised words are a few kilobytes per page.

Can I still tell which page a passage came from?

Yes. Page boundaries are marked in the output, so quotations remain attributable even though the images are not retained.

How accurate should I expect this to be?

On a clean 300 DPI printed scan in English, high — but not perfect, and errors are silent. Treat the output as a draft transcription whenever accuracy is consequential.

Read more about this