About PDF to TXT
PDF to TXT saves text decoded from a PDF’s existing text layer as a plain UTF-8 file without a byte-order mark. Choose LF or CRLF for the download; the on-page text panel keeps LF line breaks. Page markers are off by default and can be enabled when source-page references matter. This page reads all pages and groups text by position with approximate spacing, not structured columns or tables. Blank-line cleanup is on by default, reducing runs of three or more newline characters to two within each page. With cleanup off, normalized newline runs are retained and vertical gaps are used to estimate intervening blank lines, capped at 100 per gap. That is an approximation, not recovered original whitespace, and page-edge trimming still applies. Review the reading order before feeding the result to a script or importer. No OCR is performed; a document with no extractable text produces an error rather than an empty TXT download.
How to PDF to TXT
- 01
Load the PDF
Drop in one PDF. It is parsed in your browser without uploading the document for processing.
- 02
Choose the text options
Select LF or CRLF for the download, keep or disable blank-line cleanup, and enable page markers only if you need them.
- 03
Extract and review
Run the conversion and check the plain-text panel against the PDF, especially for missing characters, multi-column reading order and table alignment.
- 04
Save the .txt
Download the UTF-8 TXT file with your chosen line endings. The panel remains plain text with LF line breaks even when the download uses CRLF.
What this tool does
- UTF-8 output for the Unicode text the PDF engine can decode
- Choose LF or CRLF download line endings, independent of your operating system
- Optional blank-line cleanup, with bounded approximate vertical-gap spacing when disabled
- Optional source-page markers, off by default, in both the panel and download
- Plain text with no markup, no BOM surprises and no wrapper format
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- TXT does not preserve font styling, semantic headings, table cells or images. PDF to HTML can infer headings from text but does not restore the source layout or formatting; retain the PDF if its visual formatting is essential.
- An image-only PDF without OCR text cannot be extracted. If no page yields text, conversion stops with an error and no TXT download; use OCR PDF for scans.
- Multi-column pages flatten into a single stream, so a two-column journal article may interleave oddly.
- Spaces may approximate the positions of table columns, but there are no cell relationships or guaranteed alignment. Check tabular data before importing it.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about PDF to TXT
How is this different from the PDF to Text tool?
Both use the same extractor. PDF to Text offers page selection and a choice of stream-order or approximate positional layout, with markers on by default. PDF to TXT reads all pages using positional grouping and offers LF or CRLF downloads and optional blank-line cleanup, with markers off by default.
Can I choose CRLF line endings?
Yes. Select CRLF to write Windows-style line endings in the downloaded TXT file. LF is the default. Both are UTF-8 without a byte-order mark, and the plain-text panel continues to use LF regardless of the download setting.
What encoding is the file written in?
UTF-8, without an added byte-order mark. It encodes the Unicode text returned by the PDF engine, but cannot repair missing or incorrect font mappings. Garbled source decoding can still appear in a valid UTF-8 file.
Why was no TXT file produced?
If no page provides extractable text, the tool stops with a no-text-layer error rather than creating an empty file. Image-only scans are one cause, but blank pages or unreadable text can also trigger this result. Check the error and use OCR PDF when the visible words exist only as pixels.
Tools that pair with this one
- PDF to TextRecover the words from a PDF — read out of its text layer, not guessed.
- PDF to HTMLExtracted text in HTML page sections, with optional inferred headings.
- OCR PDFTurn a scanned PDF into something you can search — without uploading it.
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- PDF Word CounterA real word count for a PDF, with the counting rule stated up front.
- Compare PDF TextOnly the words, so a reformatted document still reads as unchanged.
Read more about this
- How to convert a PDF to JPGRendering a page to an image is a one-way conversion from instructions to pixels. This guide covers the DPI arithmetic and the JPG versus PNG decision.5 min guide
- How to convert a PDF to WordA PDF has no paragraphs to convert — only positioned glyphs. This guide explains what the converter infers, where it goes wrong, and when to extract text instead.6 min guide