From pixels to character codes
The word optical is the important one. OCR operates on an image, and its output is a sequence of character codes with positions. Between those two states sits a pipeline, and each stage is a place where quality is won or lost. First the engine cleans up: converting to greyscale, deciding a threshold that separates ink from paper, removing speckle, and correcting skew so that text baselines run horizontally. Then it segments — finding blocks of text on the page, then lines within blocks, then words, then in many engines individual glyph shapes.
Only after segmentation does recognition proper happen. Classical engines compared the isolated shape against learned feature descriptions; modern engines typically feed the line image to a neural network trained to emit a character sequence directly, which handles touching and broken letters better because it never has to commit to a hard glyph boundary. Either way the engine produces candidates with confidence scores rather than certainties.
A final stage reconciles those candidates against a language model or dictionary. This is why OCR often gets a word right despite one badly-formed letter, and also why it occasionally converts a correctly-read but unusual string — a part number, a surname, a chemical name — into a common word that looks nothing like what was on the paper. The language model is usually helping and occasionally lying.