Skip to content

7 min read · Updated 2026-09-19

How to OCR a PDF

About the recognition itself: what the engine is doing to your pixels, what raises the error rate, and which problems are fixable before you run it.

Optical character recognition turns pictures of letters into characters. It is pattern matching against trained models, which means it is probabilistic — there is always an error rate, and almost everything you can do to lower it happens before you press the button.

5 steps

The procedure

Steps for a MyPDFilles tool run in your browser tab. When a guide directs you to external software, its own privacy and security practices apply.

  1. 01

    Add the scanned PDF

    Drop in the file. The first run for a given language downloads a recognition engine and language model, which is the one network request involved — your document itself never leaves the tab.

  2. 02

    Select the document language

    Choose from English, French, Spanish, German, Italian, Portuguese or Arabic. This is not a formality: the model carries letter shapes and word patterns for that language, and the wrong choice measurably degrades accuracy on accented characters in particular.

  3. 03

    Choose what you want out

    Searchable PDF keeps the scan and adds a text layer behind it. Plain text only discards the images and gives you the recognised characters — the right choice when you want the content rather than the document.

  4. 04

    Restrict the pages if you can

    The pages field accepts all or a range. Recognition is the most computationally expensive thing on this site, so on a long document, recognising the twenty pages you need beats recognising three hundred.

  5. 05

    Run it and proofread the output

    Each page is rendered, then recognised on your CPU. Long documents take minutes. Read the result against the scan afterwards, paying attention to numbers, where an error is both likely and consequential.

Beyond the steps

What is actually happening

01

Why roughly 300 DPI is the sweet spot

Recognition works on the shape of each character, and shape needs pixels. The working rule is that a lower-case letter wants something like 20 to 30 pixels of height for a recogniser to distinguish it reliably. For ordinary body text around 10 to 12 points, that lands at approximately 300 DPI.

Below it, accuracy falls away and it falls away on specific confusions: the thin gap that separates rn from m disappears, and 8 versus B, 0 versus O, 1 versus l and comma versus period all become guesses. At 150 DPI a clean page of large type may still do well, while anything small or slightly soft degrades sharply.

Above it there is very little to gain and real cost to pay. At 600 DPI the character shapes are no better defined for the model’s purposes, while each page carries four times the pixels — four times the memory and processing time, which in a browser tab is the difference between finishing and running out of room.

If you are doing the scanning, scan at 300 DPI. If you were handed a 150 DPI scan, run it and read the output rather than upscaling: interpolating pixels invents detail and does not recover character shapes that were never captured.

02

What makes a scan hard to recognise

Skew is the most damaging and the most common. Recognisers work along text lines, so a page rotated even a couple of degrees means a line drifts across the row of pixels being analysed and characters are sliced. A scan straight in the feeder is worth more than any setting.

Low contrast is next. A faded photocopy, pencil, or a scan with the brightness pushed so that grey text sits close to grey paper leaves the engine unable to separate ink from background cleanly. High contrast — dark ink, white paper — is what the models were trained on.

Noise works against you in the same way: speckle from a dirty platen, JPEG artefacts from an over-compressed scan, show-through from the reverse of a thin page. Each adds marks that can be read as punctuation or can break a character apart. This is one reason to OCR before compressing rather than after — recognition wants the cleanest pixels available.

Then there are the genuinely hard cases. Handwriting will not work: the models are trained on print, and cursive in particular produces nothing useful. Decorative and script typefaces are unreliable for the same reason. Coloured or textured backgrounds behind text interfere with the separation of ink from paper. Very small print, footnotes and dense tabular figures push character heights below what the pixels can support.

03

Why multi-column pages come out in the wrong order

A page has no reading order recorded anywhere — it is an image, and an image is a grid of pixels with no notion of which region is a column. Reading order has to be inferred from layout, and the inference that works for the overwhelming majority of documents is to read across and then down.

On a single-column page that is exactly right. On a two-column page it is exactly wrong: the engine takes the first line of the left column, then the first line of the right column, then the second line of the left, and so on, producing text that alternates between two unrelated sentences. The characters are usually recognised correctly; it is the sequence that is broken.

Newspaper layouts, academic papers, brochures and anything with sidebars or pull quotes all hit this. Pull quotes and captions are a related problem: floating text gets spliced into the middle of the paragraph it happens to sit beside.

There are two practical responses. If you only need the document searchable, it does not matter much — each word is still in the right place on the page, so a search finds it and highlights the correct region. If you need the text as a continuous document, either accept that you will be reordering it by hand, or crop the page into columns and recognise each separately so the reading order is unambiguous.

04

Setting expectations about accuracy

A clean 300 DPI scan of ordinary printed text in a supported language, straight and high contrast, recognises well enough to be genuinely useful. A faded, skewed, low-resolution scan of small print does not, and no amount of tweaking will change that — the information is not in the pixels.

We will not quote an accuracy figure, because any single number is meaningless without the document it refers to. The same engine can be near-perfect on one page and unusable on the next, and a vendor quoting one number is describing their best-case test set.

So build proofreading into the plan for anything consequential. The errors that matter are in numbers, names and codes — exactly the content where a single wrong character changes meaning and where a human reading for sense will not notice.

While you follow this

MyPDFilles tools process files in this browser tab.

For a MyPDFilles tool, your document is read from your disk into this browser tab and the result is handed to your browser’s download mechanism. We do not make the same claim for external applications discussed in educational guides; check their own privacy and security information before using them.

Verify MyPDFilles requests in your Network panel

Questions about this task

What resolution should I scan at for OCR?

Around 300 DPI. That gives ordinary body text the character height a recogniser needs. Below 200 DPI accuracy drops noticeably; above 400 you pay in memory and time without meaningfully better recognition.

Can it read handwriting?

No. The models here are trained on printed type. Handwriting, especially cursive, produces nothing useful — that is a different class of recognition problem.

Why is my two-column text jumbled?

Because reading order is inferred from layout, and the default inference is to read across the page then down. On a two-column page that interleaves the columns line by line. Crop into columns and recognise separately if you need continuous prose.

Does OCR upload my document?

No. The recognition engine and language model are downloaded to your browser the first time you use a language, then everything runs on your CPU. Your document is not transmitted.