About scan to text
This page is about paper. Scanned documents behave differently from screenshots: they come from a flatbed or a sheet-feed scanner, they carry the specific defects of that process — slight skew from a sheet pulled crooked, speckle from dust on the glass, grey backgrounds from thin paper letting the next page through, and staple shadows down one edge — and they usually arrive as a multi-page batch. Because both scanners and scanning apps produce either images or PDFs, this tool accepts both and treats them identically, rendering PDF pages to bitmaps before recognition. The advice specific to scanning is worth following: scan at 300 DPI in greyscale rather than colour, keep the page square to the edge of the glass, and clean the glass, since a speck of dust repeated on every page is recognised as punctuation on every page. Output is editable text, page by page.
How to scan to text
- 01
Scan well
Around 300 DPI, greyscale, pages square to the glass. This is where accuracy is won or lost.
- 02
Add the scans
Drop in scanned images or a scanned PDF — both are accepted and handled the same way.
- 03
Pick the document language
Choose from the seven trained models to match the language on the paper.
- 04
Recognise and edit
Collect the text, then proofread it against the scan before you rely on it.
What this tool does
- Accepts both scanned images and scanned PDFs in one place
- Multi-page batches handled in order, with per-page results
- Greyscale and colour scans both supported
- Recognition on your own hardware, appropriate for HR files, medical records and contracts
- Seven trained languages, cached locally after first use
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- Skew is the biggest avoidable accuracy loss — a page scanned a few degrees crooked recognises noticeably worse than a straight one.
- Grey show-through from double-sided thin paper can be read as stray marks.
- Handwritten annotations and signatures on a scanned form will not be transcribed reliably.
- Multi-column pages and tables can be recognised in the wrong reading order.
How your file is handled
This tool runs inside this browser tab, but it first downloads a recognition engine and language model — static files, fetched once and then cached by your browser. Your document is never part of that request: the engine comes down to your device, and your file stays on it. You can verify this in the Network panel, where you will see the engine assets download and no upload of your document.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about scan to text
Should I scan in colour or greyscale?
Greyscale, unless colour carries meaning. Greyscale files are smaller and recognise just as well, since the engine works from contrast rather than hue.
My scans are slightly crooked. Does that matter?
Yes, more than most people expect. Line detection assumes roughly horizontal text, so a few degrees of skew costs real accuracy. Straighten before recognising if you can.
Can I feed it a scanned PDF rather than images?
Yes. Scanned PDFs are accepted and each page is rendered to a bitmap before recognition, so the result is the same as scanning to images.
Will my scanned personnel or medical records be uploaded?
No. Recognition runs in your browser on your CPU. The only network traffic is the one-time download of the engine and language model.
Tools that pair with this one
- OCR PDFTurn a scanned PDF into something you can search — without uploading it.
- OCR Image to TextPoint the recogniser at a picture and get the words it contains.
- Scan to PDFAssemble a stack of scanned page images into one properly paginated document.
- Make PDF SearchableKeep the scan exactly as it looks; make Ctrl+F start working.
- Image to Text ConverterThe alternative to retyping what you can already see on screen.
- PDF to Searchable TextWhen you want the words out of a scan, not a prettier scan.
Read more about this
- How to OCR a PDFAbout the recognition itself: what the engine is doing to your pixels, what raises the error rate, and which problems are fixable before you run it.7 min guide
- How to make a PDF searchableAbout the artefact rather than the recognition: what a text layer is, why the page looks identical afterwards, and how to verify the result is really searchable.5 min guide