Skip to content

Text · 7 min read · Updated 2026-09-19

What is OCR?

Optical character recognition is pattern matching on images of glyphs. It produces a guess, and the quality of that guess is decided long before the software runs.

A scanner produces a grid of light and dark pixels. Nothing in that grid says "this is the letter g". OCR is the process of looking at the shapes in the picture and deciding which characters they most probably represent — which means it is inference, and inference can be wrong.

What is OCR?

01

From pixels to character codes

The word optical is the important one. OCR operates on an image, and its output is a sequence of character codes with positions. Between those two states sits a pipeline, and each stage is a place where quality is won or lost. First the engine cleans up: converting to greyscale, deciding a threshold that separates ink from paper, removing speckle, and correcting skew so that text baselines run horizontally. Then it segments — finding blocks of text on the page, then lines within blocks, then words, then in many engines individual glyph shapes.

Only after segmentation does recognition proper happen. Classical engines compared the isolated shape against learned feature descriptions; modern engines typically feed the line image to a neural network trained to emit a character sequence directly, which handles touching and broken letters better because it never has to commit to a hard glyph boundary. Either way the engine produces candidates with confidence scores rather than certainties.

A final stage reconciles those candidates against a language model or dictionary. This is why OCR often gets a word right despite one badly-formed letter, and also why it occasionally converts a correctly-read but unusual string — a part number, a surname, a chemical name — into a common word that looks nothing like what was on the paper. The language model is usually helping and occasionally lying.

02

What accuracy actually depends on

Resolution comes first. The engine needs enough pixels across the height of a lowercase letter to tell shapes apart, and around 300 DPI is the practical sweet spot for ordinary body text. Below roughly 200 DPI accuracy falls off noticeably, because the distinguishing features of similar glyphs stop being represented at all. Scanning far above 300 DPI mostly buys compute time and file size rather than accuracy — the exception being genuinely small print, where more pixels per character still help.

Contrast matters nearly as much. The engine has to separate ink from background, so grey text on grey paper, faded thermal receipts, photocopies of photocopies and pages shot in uneven light all cause thresholding to eat parts of letters or fuse neighbouring ones. Noise works the same way: dust, scanner streaks, JPEG compression artefacts around letterforms and show-through from the reverse of thin paper all add marks the segmenter has to interpret.

Skew is worth singling out because it is so easy to fix and so damaging when unfixed. Line-finding assumes roughly horizontal baselines; a page fed in at an angle, or photographed at a tilt with perspective distortion, can break line segmentation before recognition even begins. Typeface plays a part too: clean serif and sans-serif body text is what engines are trained on, while decorative display faces, condensed type, heavy italics, blackletter and stylised logos are materially harder. Very small type and tight letter spacing increase the chance that adjacent characters merge into one blob.

03

Where OCR should not be trusted

Handwriting is the clearest case. Recognising handwriting is a different problem from recognising print — letterforms vary within a single writer, cursive joins characters continuously, and there is no fixed alphabet of shapes to match. Some engines attempt it and some do so reasonably on neat block capitals in known fields, but treating handwritten OCR output as reliable text is a mistake.

Layout is the subtler trap. An engine reads shapes; deciding what order those shapes should be read in is a separate inference. A two-column page can be read straight across, producing sentences that alternate between columns and make no sense. Tables are harder still: cell boundaries may be implied by whitespace rather than drawn, and a value can end up attached to the wrong row or column, which is worse than an obvious error because the result still looks plausible.

Then there are the characters that are genuinely ambiguous in isolation, where the shapes are near-identical in many typefaces: the digit zero against a capital O, the digit one against a lowercase l and a capital I, a lowercase rn reading as an m, an 8 against a B. In running prose the language model resolves these. In reference numbers, serial codes, postcodes and account identifiers there is no linguistic context to resolve them, which is exactly where the errors hurt most and exactly why OCR output over such data needs checking against the image.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

What resolution should I scan at for OCR?

Around 300 DPI for ordinary body text. That gives the engine enough pixels per character to distinguish similar shapes without inflating processing time. Go higher only for very small print or fine footnotes, and avoid dropping much below 200 DPI, where accuracy degrades quickly.

Can OCR read handwriting?

Not reliably. Handwriting recognition is a distinct and much harder problem: letterforms vary within a single writer, cursive joins characters together, and there is no fixed set of shapes to match against. Neat block capitals in structured fields sometimes work, but output should never be trusted unchecked.

Why does OCR mangle tables and multi-column pages?

Because recognising characters and determining reading order are separate problems. Column and cell boundaries are often implied by whitespace rather than drawn, so the engine can read straight across two columns or attach a value to the wrong row — and the result still looks plausible, which makes it harder to catch.

Does converting a scan to greyscale help OCR?

Usually yes, because it removes colour noise before the engine decides its ink-versus-paper threshold. Most engines convert internally anyway, so the real gains come from good contrast, even lighting and a straight page rather than from the colour mode you chose.