About batch extract text from PDF
Text extraction reads the text-showing operations in each page’s content stream and reassembles the characters they draw. Done across a document set it is the practical way to make a pile of PDFs searchable, feed them to an indexer, or run any kind of analysis over their contents. The critical thing to know is that extraction only finds text that exists as text. A scanned document contains a photograph of words, with no characters in the content stream at all, so it yields an empty file — which is why the empty-file report matters: it tells you which inputs are scans needing OCR rather than leaving you to wonder why a document came back blank. Layout handling is the other real decision. Reading order produces continuous prose suited to indexing and analysis, while preserving layout keeps approximate positions, which is what you need when a page’s meaning depends on its table structure rather than on its sentences.
How to batch extract text from PDF
- 01
Add the PDFs
Drop in every document you want the text from.
- 02
Choose layout handling
Reading order for prose and indexing; preserve layout when tables matter.
- 03
Keep page markers on
They let you trace any extracted passage back to the page it came from.
- 04
Download and check the report
A .txt per document arrives in a ZIP, with any empty results flagged as probable scans.
What this tool does
- One plain .txt file per source document, named to match
- Three layout modes: reading order, preserved layout and raw stream order
- Optional page markers for tracing an extract back to its origin
- Explicit reporting of files that yielded no text, identifying scans that need OCR
- Fast even on large sets, since text extraction does not require rendering pages
Limitations worth knowing
Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.
- Scanned documents contain images, not text, and produce empty output. Run OCR PDF on those first — extraction cannot invent characters that are not there.
- Reading order follows the content stream, so multi-column pages can interleave columns into jumbled prose.
- Tables lose their structure in reading-order mode; cell boundaries are not preserved as data.
- Ligatures, unusual encodings and custom font mappings can extract as incorrect or missing characters.
How your file is handled
This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.
Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.
Questions about batch extract text from PDF
Why did some files come back empty?
They are scans — images of pages, with no characters in the content stream to extract. The report flags them precisely so you know to run OCR PDF on those rather than assuming something went wrong.
Which layout mode should I use?
Reading order for prose you will index, search or analyse. Preserve layout when the page’s meaning depends on its spatial arrangement, such as tables or forms. Raw mode is mainly a diagnostic view.
Why is the text from my two-column document jumbled?
Extraction follows the order text was written into the content stream, which in some multi-column layouts alternates across columns line by line. Preserve-layout mode usually reads better for those pages.
Can I get the text as one combined file?
No — you get one .txt per source document, which keeps results traceable. Concatenate them afterwards if you need a single corpus.
Tools that pair with this one
- Extract Text from PDFLift the text layer out of a PDF and keep it in reading order.
- OCR PDFTurn a scanned PDF into something you can search — without uploading it.
- Batch Compress PDFCompress a whole folder of PDFs in one pass and get a ZIP back.
- Batch PDF to ImagesRender every page of every file to an image, organised per document.
- Compare PDF FilesTwo versions in, a clear list of differences out.
- PDF InspectorA complete technical report on any PDF, produced without uploading it.