Skip to content

Extract From PDF

Extract Text from PDF

Lift the text layer out of a PDF and keep it in reading order.

Processed locally in your browser. Your file is not uploaded to any server. How this works

  1. 01Add your file
  2. 02Extract Text from PDF
  3. 03Download

Runs in your browser · nothing uploaded

About extract text from PDF

A PDF does not store paragraphs. It stores a sequence of drawing instructions that place short runs of glyphs at exact coordinates, which is why naive copy-paste out of a PDF so often returns jumbled fragments. This tool reads the page content streams, collects every text run with its position and font size, then reconstructs reading order geometrically: runs sharing a baseline become a line, lines sharing a horizontal band become a block, and blocks are ordered top-to-bottom within each detected column. With layout preservation on you get something close to the page as a human reads it, including indentation and column separation. With it off, runs are joined into continuous flowing prose, which is usually what you want before feeding text to another program. Everything happens in this tab against the file in memory.

How to extract text from PDF

  1. 01

    Add the PDF

    Drop the file onto the drop zone. It is parsed locally; no request carries its bytes.

  2. 02

    Pick pages

    Leave Pages set to all, or narrow it to the ranges you actually need — handy for a single chapter of a long report.

  3. 03

    Choose a layout mode

    Keep Preserve layout on to mirror the page visually, or switch it off to get continuous prose.

  4. 04

    Copy or download

    Read the result in the preview, copy it to the clipboard, or save it as a .txt file.

What this tool does

  • Geometric reading-order reconstruction rather than raw content-stream order
  • Page-range selection so you can extract one section of a long document
  • Two output shapes: layout-faithful, or continuous prose for further processing
  • Ligatures and common glyph substitutions normalised back to plain characters
  • Runs entirely in the tab — no upload, no retention, works offline after first load

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • Scanned pages hold pictures of words, not text. They return nothing here — run OCR PDF first.
  • Tables come out as lines of text; cell boundaries are not reconstructed into columns or CSV.
  • A few PDFs embed fonts with broken character maps, so extracted characters can be wrong even though the page looks fine.
  • Text drawn inside vector artwork or as outlines is not text to the file format and cannot be extracted.

How your file is handled

This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about extract text from PDF

Why does copy-paste from a PDF reader scramble text but this tool does not?

Readers often hand back runs in the order they were drawn. This tool sorts runs by their coordinates first, rebuilding lines and columns, so the order matches how the page reads.

Nothing was extracted from my file. What happened?

Almost certainly a scan. If you can select a word in a PDF reader there is a text layer; if you cannot, the page is an image and needs OCR before any text exists to extract.

Should I leave layout preservation on?

Leave it on when you want the text to look like the page — reports, letters, anything with columns. Turn it off when the text is going into another tool and paragraph continuity matters more than appearance.

Do you see the text that comes out?

No. Parsing and reconstruction run in JavaScript inside your tab, and the result exists only in that tab until you copy or download it.

Read more about this