About PDF to text
PDF to Text reads the text layer already stored in a PDF; it does not recognise words in images. By default, it follows the browser PDF engine’s text-stream order and line-break signals. Enable Approximate layout to group text by position and add approximate spaces and line breaks instead. Either mode can change whitespace, interleave columns or misread characters when font mappings are unreliable, so the result is not a character-for-character copy of the author’s source. Choose all pages or a page selection, then copy the plain-text panel or download a UTF-8 TXT file. Page markers are enabled by default and use original source-page numbers; turn them off to omit generated labels from both outputs. Literal text that happens to resemble a marker is not removed. If none of the selected pages yields text, the run stops with an error instead of producing an empty download. Image-only pages need OCR.
How to PDF to text
- 01
Add the PDF
Drop the file in. It is parsed locally by an in-browser PDF engine.
- 02
Choose pages, layout and markers
Keep all or enter a selection such as 1-10. Leave layout off for stream order, or enable approximate positional grouping. Keep page markers when source-page references are useful.
- 03
Extract
The text layer is walked page by page and the characters are decoded through each font's encoding.
- 04
Copy or download
Copy the plain text from the panel or download the UTF-8 TXT file. The PDF and extracted text are not uploaded for processing.
What this tool does
- Reads available text from the PDF’s text layer rather than recognising pixels
- Choose PDF-engine text-stream order or position-based grouping with approximate spacing
- Optional page markers identify selected pages by their original source numbers
- Page ranges, so you can read a chapter out of a long document
- UTF-8 output for decoded text; character accuracy depends on the PDF and its font mappings
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- If no selected page has extractable text, the tool reports an error and produces no download. Image-only scans need OCR first; blank pages and unreadable text mappings can also yield no text.
- Line breaks come from PDF-engine signals or positional grouping, depending on the layout setting. Neither uses PDF structure tags to recover semantic reading order; columns and inferred breaks can be wrong.
- Missing or unusual font mappings and decoder limitations can produce missing or garbled characters. Check important passages against the PDF rather than assuming extraction is exact.
- Text drawn as vector outlines rather than as glyphs — common in logos and some design exports — is invisible to extraction.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about PDF to text
Why did I get no text at all?
No extractable text was found on the selected pages, so no file was produced. They may be image-only scans, blank pages or text drawn as outlines; decoding problems can also prevent extraction. Check the page selection and use OCR PDF when the words exist only as an image.
How do I tell whether my PDF has a text layer?
Try selecting and copying a line in a PDF viewer. Readable copied words are a useful indication of an extractable text layer, but viewer selection behavior is not a definitive test. A scan can already have OCR text behind its image, and a text layer can still decode poorly.
Why is the text in a strange order on some pages?
With Approximate layout off, extraction follows the PDF engine’s text-stream order, which may not match reading order. With it on, lines are grouped from top to bottom and runs from left to right. Neither mode identifies semantic columns, so multi-column pages can interleave. Compare both settings and review important passages.
Is the extracted text exactly what the author typed?
Not necessarily. The tool decodes stored text rather than recognising pixels, but it also reorders runs, inserts spacing and normalizes whitespace. Font mappings and PDF decoding can introduce missing or incorrect characters, so verify text that needs exact fidelity.
Does anything get uploaded?
Processing does not upload your PDF or extracted text. The browser may fetch application and engine files; consented usage events can be sent without document contents.
Tools that pair with this one
- PDF to TXTChoose line endings and page markers for a UTF-8 text download.
- PDF to HTMLExtracted text in HTML page sections, with optional inferred headings.
- OCR PDFTurn a scanned PDF into something you can search — without uploading it.
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- PDF Word CounterA real word count for a PDF, with the counting rule stated up front.
- PDF to WordAn editable Word document rebuilt from your PDF’s text and positions.
Read more about this
- How to convert a PDF to WordA PDF has no paragraphs to convert — only positioned glyphs. This guide explains what the converter infers, where it goes wrong, and when to extract text instead.6 min guide
- How to convert a PDF to JPGRendering a page to an image is a one-way conversion from instructions to pixels. This guide covers the DPI arithmetic and the JPG versus PNG decision.5 min guide