About OCR PDF
A scanned PDF is a photo album: each page is an image, and there is no text in the file at all, which is why search finds nothing and selection grabs nothing. OCR fixes that by recognising the shapes. Each page is rendered to a bitmap, then Tesseract — compiled to WebAssembly and running on your own CPU — segments the image into lines and characters, classifies each glyph against a trained model for the language you chose, and uses that language's word patterns to resolve ambiguous shapes. In searchable-PDF mode the recognised words are written back as an invisible text layer positioned exactly behind the words in the image: the page looks pixel-identical to the original scan, but text can be selected, searched and copied, because your reader finds the hidden layer underneath. Because the work is local it is slower than a server — expect a few seconds per page — and the first run downloads a multi-megabyte engine and language model, which the browser then caches.
How to OCR PDF
- 01
Add the scanned PDF
Drop the file in. It stays on your device; only the OCR engine is fetched from the network.
- 02
Set the language
Choose the language the document is written in. This matters more for accuracy than any other setting.
- 03
Choose the output
Searchable PDF keeps the original page appearance and adds a hidden text layer. Text-only returns just the words.
- 04
Run it and wait
Recognition runs page by page on your CPU. A first run also downloads the engine and model, then caches them.
What this tool does
- Tesseract running locally via WebAssembly — the scan is never uploaded
- Seven trained languages: English, French, Spanish, German, Italian, Portuguese and Arabic
- Searchable-PDF output keeps the original image and hides the text behind it
- Page ranges, so you can OCR a chapter instead of a 300-page book
- Engine and language model cached by the browser after the first run
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- Only the seven languages listed are available; there is no model here for Chinese, Japanese, Korean, Russian, Hindi or others.
- Handwriting is not reliably recognised. These models are trained on printed type.
- Accuracy depends on the scan: around 300 DPI, good contrast and straight pages give the best results, while low-resolution or skewed scans degrade sharply.
- Complex multi-column layouts and tables can be read in the wrong order, because reading order is inferred from the image rather than known.
- Local OCR is slower than a server and uses your CPU; a long document takes minutes, not seconds.
How your file is handled
This tool runs inside this browser tab, but it first downloads a recognition engine and language model — static files, fetched once and then cached by your browser. Your document is never part of that request: the engine comes down to your device, and your file stays on it. You can verify this in the Network panel, where you will see the engine assets download and no upload of your document.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about OCR PDF
What exactly is a "searchable PDF"?
Your original scanned image, unchanged, with an invisible text layer placed behind it at the same coordinates as the words in the picture. The page looks identical, but selecting or searching hits the hidden text underneath.
Why is it slower than an online OCR service?
Because the recognition is happening on your processor rather than on a datacentre machine. That is the trade for the file never being uploaded. Budget a few seconds per page.
Why does the first run download something?
The WebAssembly build of Tesseract and the trained data for your chosen language are several megabytes and must be present locally to run. They are fetched once and then cached by your browser.
How good is Arabic recognition?
Usable but genuinely harder than the Latin-script languages. Arabic is right-to-left and cursive, with letters changing shape by position, so expect more errors than English on a comparable scan.
What resolution should I scan at?
About 300 DPI. Lower loses the detail that distinguishes similar letters; much higher mostly adds processing time without improving accuracy.
Tools that pair with this one
- Make PDF SearchableKeep the scan exactly as it looks; make Ctrl+F start working.
- OCR Image to TextPoint the recogniser at a picture and get the words it contains.
- PDF to Searchable TextWhen you want the words out of a scan, not a prettier scan.
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- Scan to PDFAssemble a stack of scanned page images into one properly paginated document.
- Compress PDFMake a PDF smaller by re-encoding its images and cleaning up its internals.
Read more about this
- How to OCR a PDFAbout the recognition itself: what the engine is doing to your pixels, what raises the error rate, and which problems are fixable before you run it.7 min guide
- How to make a PDF searchableAbout the artefact rather than the recognition: what a text layer is, why the page looks identical afterwards, and how to verify the result is really searchable.5 min guide