Skip to content

Convert From PDF

PDF to TXT

Choose line endings and page markers for a UTF-8 text download.

Processed locally in your browser. Your file is not uploaded to any server. How this works

  1. 01Add your file
  2. 02PDF to TXT
  3. 03Download

Runs in your browser · nothing uploaded

About PDF to TXT

PDF to TXT saves text decoded from a PDF’s existing text layer as a plain UTF-8 file without a byte-order mark. Choose LF or CRLF for the download; the on-page text panel keeps LF line breaks. Page markers are off by default and can be enabled when source-page references matter. This page reads all pages and groups text by position with approximate spacing, not structured columns or tables. Blank-line cleanup is on by default, reducing runs of three or more newline characters to two within each page. With cleanup off, normalized newline runs are retained and vertical gaps are used to estimate intervening blank lines, capped at 100 per gap. That is an approximation, not recovered original whitespace, and page-edge trimming still applies. Review the reading order before feeding the result to a script or importer. No OCR is performed; a document with no extractable text produces an error rather than an empty TXT download.

How to PDF to TXT

  1. 01

    Load the PDF

    Drop in one PDF. It is parsed in your browser without uploading the document for processing.

  2. 02

    Choose the text options

    Select LF or CRLF for the download, keep or disable blank-line cleanup, and enable page markers only if you need them.

  3. 03

    Extract and review

    Run the conversion and check the plain-text panel against the PDF, especially for missing characters, multi-column reading order and table alignment.

  4. 04

    Save the .txt

    Download the UTF-8 TXT file with your chosen line endings. The panel remains plain text with LF line breaks even when the download uses CRLF.

What this tool does

  • UTF-8 output for the Unicode text the PDF engine can decode
  • Choose LF or CRLF download line endings, independent of your operating system
  • Optional blank-line cleanup, with bounded approximate vertical-gap spacing when disabled
  • Optional source-page markers, off by default, in both the panel and download
  • Plain text with no markup, no BOM surprises and no wrapper format

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • TXT does not preserve font styling, semantic headings, table cells or images. PDF to HTML can infer headings from text but does not restore the source layout or formatting; retain the PDF if its visual formatting is essential.
  • An image-only PDF without OCR text cannot be extracted. If no page yields text, conversion stops with an error and no TXT download; use OCR PDF for scans.
  • Multi-column pages flatten into a single stream, so a two-column journal article may interleave oddly.
  • Spaces may approximate the positions of table columns, but there are no cell relationships or guaranteed alignment. Check tabular data before importing it.

How your file is handled

This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about PDF to TXT

How is this different from the PDF to Text tool?

Both use the same extractor. PDF to Text offers page selection and a choice of stream-order or approximate positional layout, with markers on by default. PDF to TXT reads all pages using positional grouping and offers LF or CRLF downloads and optional blank-line cleanup, with markers off by default.

Can I choose CRLF line endings?

Yes. Select CRLF to write Windows-style line endings in the downloaded TXT file. LF is the default. Both are UTF-8 without a byte-order mark, and the plain-text panel continues to use LF regardless of the download setting.

What encoding is the file written in?

UTF-8, without an added byte-order mark. It encodes the Unicode text returned by the PDF engine, but cannot repair missing or incorrect font mappings. Garbled source decoding can still appear in a valid UTF-8 file.

Why was no TXT file produced?

If no page provides extractable text, the tool stops with a no-text-layer error rather than creating an empty file. Image-only scans are one cause, but blank pages or unreadable text can also trigger this result. Check the error and use OCR PDF when the visible words exist only as pixels.

Read more about this

Background