Skip to content

Text · 6 min read · Updated 2026-09-19

What is a searchable PDF?

A two-layer artefact: the original page image you see, plus invisible text underneath it that search and selection actually operate on.

A searchable PDF is not a PDF whose scan has been turned into text. It is a PDF that carries both things at once — the untouched page picture and a hidden transcript positioned to match it. Understanding that layering explains both why it is so useful and where it quietly misleads.

What is a searchable PDF?

01

Two layers, one page

Start from what a scan gives you: one raster image per page, wrapped in a PDF. It looks like a document and behaves like a photograph. Select a word and you select nothing. Search for a phrase and you find nothing, because no character codes exist anywhere in the file.

Making it searchable does not replace that image. The image stays exactly as it was, as the visible content of the page. What is added is a second set of content: text-showing operations placing the recognised characters at the coordinates where those characters appear in the picture, word by word and often glyph by glyph. The page now contains a drawn image and a full text transcript occupying the same rectangle.

The transcript is made invisible by PDF text render mode 3, which means neither fill nor stroke — the glyphs are positioned and measured, and then nothing is painted. They are real text as far as the file is concerned. They are simply not rendered. Some producers additionally place the text behind the image in drawing order, so even a rendering quirk would leave it covered.

The consequence is that the page is pixel-identical to the original scan while being fully searchable. Searching walks the text layer. Selecting a sentence hits the invisible glyph boxes and highlights the region of the image they sit over, which is why a selection rectangle in a scanned PDF can look slightly offset from the printed ink — you are selecting the transcript, not the picture.

02

The consequence nobody mentions: invisible errors

Because the visible layer is untouched, recognition mistakes are invisible. This deserves stating bluntly, since it is the single most important property of the artefact. If the hidden text says Invoice 1O834 where the page shows Invoice 10834, the page still displays the correct number in the original ink. Nothing on screen looks wrong. Nothing can look wrong, because what you are looking at was never modified.

The error only surfaces when something consumes the text layer rather than the image. A search for the correct invoice number returns nothing. A copy-paste yields the wrong string. A full-text index built across an archive silently omits the document from results it should have matched. A script extracting values gets a mis-transcribed figure and carries on. The document looks perfect and behaves incorrectly, which is a harder failure to notice than a visibly garbled page.

This cuts the other way too, and it is the reason the design is sound. Because the image is preserved, an OCR error costs you a search hit rather than the content itself. The authoritative version of the page is still there, exactly as scanned, available to read and to re-recognise later with a better engine. Contrast that with replacing the page with recognised text, where a mistake becomes the document and the evidence of it is gone.

Practically, that means treating the text layer as an index and the image as the record. Where the extracted values matter — amounts, identifiers, dates, names — they should be checked against the visible page rather than trusted because they came out of a file that looked fine.

03

What searchable PDFs are not

They are not editable documents. The visible page is an image, so there is no text to revise — changing what the page says would mean altering pixels. Adding a hidden transcript gives you search and copy, not editing.

They are not small. The page image dominates the file size, and the text layer adds very little on top. Making a scan searchable therefore barely changes its weight; if a scanned PDF is uncomfortably large, the answer is image resolution and compression, not the text layer.

They are not automatically accessible. A hidden text layer gives a screen reader something to read, which is a real improvement over a bare image, but it carries no logical structure — no headings, no reading order guarantees beyond the order the recognition engine emitted, no table relationships. Genuine accessibility needs tagging on top.

And they are not guaranteed to be in reading order. The sequence of the invisible text follows whatever order recognition produced, so a two-column scan can yield a transcript that alternates between columns even though the page looks perfectly normal. Extraction from such a file returns interleaved lines with no visible clue that anything went wrong.

Worth repeating

MyPDFilles tools described here run on your device.

When an article refers to a MyPDFilles tool, its parsing, compression or recognition runs in JavaScript and WebAssembly inside your browser tab on bytes read from your disk. Educational references to external software are not covered by that claim; review the external provider’s own privacy and security information.

How to verify MyPDFilles processing

Questions on this topic

How can text be in a PDF but invisible?

PDF has a text rendering mode that neither fills nor strokes the glyphs — mode 3. The characters are positioned, sized and fully present in the content stream, but nothing is painted for them. Search, selection and extraction all work on them because they are real text; your eyes simply never receive anything.

Will making a scan searchable change how the page looks?

No. The original page image is kept exactly as it was and remains the only visible content. The recognised text is added as a separate invisible layer aligned to it, so the rendered page is pixel-identical to the scan you started with.

How do I know whether the hidden text is correct?

You have to check it deliberately, because the page cannot show you. Copy text out or extract it and compare against the visible image, paying particular attention to identifiers, amounts and dates where no dictionary can correct a misread character.

Does a searchable PDF get much bigger?

Barely. Character codes and positions are tiny compared with a page image, so the added layer is a small fraction of the file. The size of a scanned PDF is driven almost entirely by the resolution and compression of its images.