Two layers, one page
Start from what a scan gives you: one raster image per page, wrapped in a PDF. It looks like a document and behaves like a photograph. Select a word and you select nothing. Search for a phrase and you find nothing, because no character codes exist anywhere in the file.
Making it searchable does not replace that image. The image stays exactly as it was, as the visible content of the page. What is added is a second set of content: text-showing operations placing the recognised characters at the coordinates where those characters appear in the picture, word by word and often glyph by glyph. The page now contains a drawn image and a full text transcript occupying the same rectangle.
The transcript is made invisible by PDF text render mode 3, which means neither fill nor stroke — the glyphs are positioned and measured, and then nothing is painted. They are real text as far as the file is concerned. They are simply not rendered. Some producers additionally place the text behind the image in drawing order, so even a rendering quirk would leave it covered.
The consequence is that the page is pixel-identical to the original scan while being fully searchable. Searching walks the text layer. Selecting a sentence hits the invisible glyph boxes and highlights the region of the image they sit over, which is why a selection rectangle in a scanned PDF can look slightly offset from the printed ink — you are selecting the transcript, not the picture.