Skip to content

PDF OCR

OCR PDF

Turn a scanned PDF into something you can search — without uploading it.

Processed locally in your browser. Downloads a recognition engine once; your file is still never uploaded. How this works

  1. 01Add your file
  2. 02OCR PDF
  3. 03Download

Runs in your browser · downloads an engine file once

About OCR PDF

A scanned PDF is a photo album: each page is an image, and there is no text in the file at all, which is why search finds nothing and selection grabs nothing. OCR fixes that by recognising the shapes. Each page is rendered to a bitmap, then Tesseract — compiled to WebAssembly and running on your own CPU — segments the image into lines and characters, classifies each glyph against a trained model for the language you chose, and uses that language's word patterns to resolve ambiguous shapes. In searchable-PDF mode the recognised words are written back as an invisible text layer positioned exactly behind the words in the image: the page looks pixel-identical to the original scan, but text can be selected, searched and copied, because your reader finds the hidden layer underneath. Because the work is local it is slower than a server — expect a few seconds per page — and the first run downloads a multi-megabyte engine and language model, which the browser then caches.

How to OCR PDF

  1. 01

    Add the scanned PDF

    Drop the file in. It stays on your device; only the OCR engine is fetched from the network.

  2. 02

    Set the language

    Choose the language the document is written in. This matters more for accuracy than any other setting.

  3. 03

    Choose the output

    Searchable PDF keeps the original page appearance and adds a hidden text layer. Text-only returns just the words.

  4. 04

    Run it and wait

    Recognition runs page by page on your CPU. A first run also downloads the engine and model, then caches them.

What this tool does

  • Tesseract running locally via WebAssembly — the scan is never uploaded
  • Seven trained languages: English, French, Spanish, German, Italian, Portuguese and Arabic
  • Searchable-PDF output keeps the original image and hides the text behind it
  • Page ranges, so you can OCR a chapter instead of a 300-page book
  • Engine and language model cached by the browser after the first run

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • Only the seven languages listed are available; there is no model here for Chinese, Japanese, Korean, Russian, Hindi or others.
  • Handwriting is not reliably recognised. These models are trained on printed type.
  • Accuracy depends on the scan: around 300 DPI, good contrast and straight pages give the best results, while low-resolution or skewed scans degrade sharply.
  • Complex multi-column layouts and tables can be read in the wrong order, because reading order is inferred from the image rather than known.
  • Local OCR is slower than a server and uses your CPU; a long document takes minutes, not seconds.

How your file is handled

This tool runs inside this browser tab, but it first downloads a recognition engine and language model — static files, fetched once and then cached by your browser. Your document is never part of that request: the engine comes down to your device, and your file stays on it. You can verify this in the Network panel, where you will see the engine assets download and no upload of your document.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about OCR PDF

What exactly is a "searchable PDF"?

Your original scanned image, unchanged, with an invisible text layer placed behind it at the same coordinates as the words in the picture. The page looks identical, but selecting or searching hits the hidden text underneath.

Why is it slower than an online OCR service?

Because the recognition is happening on your processor rather than on a datacentre machine. That is the trade for the file never being uploaded. Budget a few seconds per page.

Why does the first run download something?

The WebAssembly build of Tesseract and the trained data for your chosen language are several megabytes and must be present locally to run. They are fetched once and then cached by your browser.

How good is Arabic recognition?

Usable but genuinely harder than the Latin-script languages. Arabic is right-to-left and cursive, with letters changing shape by position, so expect more errors than English on a comparable scan.

What resolution should I scan at?

About 300 DPI. Lower loses the detail that distinguishes similar letters; much higher mostly adds processing time without improving accuracy.

Read more about this