Skip to content

Convert From PDF

PDF to Text

Recover the words from a PDF — read out of its text layer, not guessed.

Processed locally in your browser. Your file is not uploaded to any server. How this works

  1. 01Add your file
  2. 02PDF to Text
  3. 03Download

Runs in your browser · nothing uploaded

About PDF to text

PDF to Text reads the text layer already stored in a PDF; it does not recognise words in images. By default, it follows the browser PDF engine’s text-stream order and line-break signals. Enable Approximate layout to group text by position and add approximate spaces and line breaks instead. Either mode can change whitespace, interleave columns or misread characters when font mappings are unreliable, so the result is not a character-for-character copy of the author’s source. Choose all pages or a page selection, then copy the plain-text panel or download a UTF-8 TXT file. Page markers are enabled by default and use original source-page numbers; turn them off to omit generated labels from both outputs. Literal text that happens to resemble a marker is not removed. If none of the selected pages yields text, the run stops with an error instead of producing an empty download. Image-only pages need OCR.

How to PDF to text

  1. 01

    Add the PDF

    Drop the file in. It is parsed locally by an in-browser PDF engine.

  2. 02

    Choose pages, layout and markers

    Keep all or enter a selection such as 1-10. Leave layout off for stream order, or enable approximate positional grouping. Keep page markers when source-page references are useful.

  3. 03

    Extract

    The text layer is walked page by page and the characters are decoded through each font's encoding.

  4. 04

    Copy or download

    Copy the plain text from the panel or download the UTF-8 TXT file. The PDF and extracted text are not uploaded for processing.

What this tool does

  • Reads available text from the PDF’s text layer rather than recognising pixels
  • Choose PDF-engine text-stream order or position-based grouping with approximate spacing
  • Optional page markers identify selected pages by their original source numbers
  • Page ranges, so you can read a chapter out of a long document
  • UTF-8 output for decoded text; character accuracy depends on the PDF and its font mappings

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • If no selected page has extractable text, the tool reports an error and produces no download. Image-only scans need OCR first; blank pages and unreadable text mappings can also yield no text.
  • Line breaks come from PDF-engine signals or positional grouping, depending on the layout setting. Neither uses PDF structure tags to recover semantic reading order; columns and inferred breaks can be wrong.
  • Missing or unusual font mappings and decoder limitations can produce missing or garbled characters. Check important passages against the PDF rather than assuming extraction is exact.
  • Text drawn as vector outlines rather than as glyphs — common in logos and some design exports — is invisible to extraction.

How your file is handled

This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about PDF to text

Why did I get no text at all?

No extractable text was found on the selected pages, so no file was produced. They may be image-only scans, blank pages or text drawn as outlines; decoding problems can also prevent extraction. Check the page selection and use OCR PDF when the words exist only as an image.

How do I tell whether my PDF has a text layer?

Try selecting and copying a line in a PDF viewer. Readable copied words are a useful indication of an extractable text layer, but viewer selection behavior is not a definitive test. A scan can already have OCR text behind its image, and a text layer can still decode poorly.

Why is the text in a strange order on some pages?

With Approximate layout off, extraction follows the PDF engine’s text-stream order, which may not match reading order. With it on, lines are grouped from top to bottom and runs from left to right. Neither mode identifies semantic columns, so multi-column pages can interleave. Compare both settings and review important passages.

Is the extracted text exactly what the author typed?

Not necessarily. The tool decodes stored text rather than recognising pixels, but it also reorders runs, inserts spacing and normalizes whitespace. Font mappings and PDF decoding can introduce missing or incorrect characters, so verify text that needs exact fidelity.

Does anything get uploaded?

Processing does not upload your PDF or extracted text. The browser may fetch application and engine files; consented usage events can be sent without document contents.

Read more about this