Skip to content

PDF OCR

Scan to Text

For paper that has been scanned and now needs to be text.

Processed locally in your browser. Downloads a recognition engine once; your file is still never uploaded. How this works

  1. 01Add your files
  2. 02Scan to Text
  3. 03Download

Runs in your browser · downloads an engine file once

About scan to text

This page is about paper. Scanned documents behave differently from screenshots: they come from a flatbed or a sheet-feed scanner, they carry the specific defects of that process — slight skew from a sheet pulled crooked, speckle from dust on the glass, grey backgrounds from thin paper letting the next page through, and staple shadows down one edge — and they usually arrive as a multi-page batch. Because both scanners and scanning apps produce either images or PDFs, this tool accepts both and treats them identically, rendering PDF pages to bitmaps before recognition. The advice specific to scanning is worth following: scan at 300 DPI in greyscale rather than colour, keep the page square to the edge of the glass, and clean the glass, since a speck of dust repeated on every page is recognised as punctuation on every page. Output is editable text, page by page.

How to scan to text

  1. 01

    Scan well

    Around 300 DPI, greyscale, pages square to the glass. This is where accuracy is won or lost.

  2. 02

    Add the scans

    Drop in scanned images or a scanned PDF — both are accepted and handled the same way.

  3. 03

    Pick the document language

    Choose from the seven trained models to match the language on the paper.

  4. 04

    Recognise and edit

    Collect the text, then proofread it against the scan before you rely on it.

What this tool does

  • Accepts both scanned images and scanned PDFs in one place
  • Multi-page batches handled in order, with per-page results
  • Greyscale and colour scans both supported
  • Recognition on your own hardware, appropriate for HR files, medical records and contracts
  • Seven trained languages, cached locally after first use

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • Skew is the biggest avoidable accuracy loss — a page scanned a few degrees crooked recognises noticeably worse than a straight one.
  • Grey show-through from double-sided thin paper can be read as stray marks.
  • Handwritten annotations and signatures on a scanned form will not be transcribed reliably.
  • Multi-column pages and tables can be recognised in the wrong reading order.

How your file is handled

This tool runs inside this browser tab, but it first downloads a recognition engine and language model — static files, fetched once and then cached by your browser. Your document is never part of that request: the engine comes down to your device, and your file stays on it. You can verify this in the Network panel, where you will see the engine assets download and no upload of your document.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about scan to text

Should I scan in colour or greyscale?

Greyscale, unless colour carries meaning. Greyscale files are smaller and recognise just as well, since the engine works from contrast rather than hue.

My scans are slightly crooked. Does that matter?

Yes, more than most people expect. Line detection assumes roughly horizontal text, so a few degrees of skew costs real accuracy. Straighten before recognising if you can.

Can I feed it a scanned PDF rather than images?

Yes. Scanned PDFs are accepted and each page is rendered to a bitmap before recognition, so the result is the same as scanning to images.

Will my scanned personnel or medical records be uploaded?

No. Recognition runs in your browser on your CPU. The only network traffic is the one-time download of the engine and language model.

Read more about this