Skip to content

6 min read · Updated 2026-09-19

How to convert a PDF to Word

A PDF has no paragraphs to convert — only positioned glyphs. This guide explains what the converter infers, where it goes wrong, and when to extract text instead.

This is the conversion with the widest gap between expectation and reality, and the reason is structural rather than a matter of tool quality. A Word document describes a flowing sequence of paragraphs; a PDF describes marks at coordinates. Going from the second to the first means inferring a structure that was never recorded.

5 steps

The procedure

Steps for a MyPDFilles tool run in your browser tab. When a guide directs you to external software, its own privacy and security practices apply.

  1. 01

    Add the PDF

    Drop in the file. Before going further, click into a page in any viewer and try to select a line of text. If you cannot, the pages are images and there is nothing for this tool to read — run OCR first.

  2. 02

    Decide about heading styles

    Map headings to Word heading styles is on by default. It infers headings from relative font size and emits them as Heading styles, which gives you a navigable document instead of a wall of body text.

  3. 03

    Decide about bold and italic

    Keep bold and italic is also on by default. It reads the font name of each glyph run and applies character formatting where the run used a bold or italic face.

  4. 04

    Restrict the pages if you only need part

    The pages field takes all or a list such as 1, 4, 7-9. Converting only the section you intend to edit produces a cleaner document than converting a hundred pages and deleting most of them.

  5. 05

    Convert, then repair in Word

    Run it and open the result. Expect to spend a few minutes fixing paragraph breaks and any table that came out as loose lines — that clean-up is inherent to the operation, not a sign something failed.

Beyond the steps

What is actually happening

01

Why this is a reconstruction rather than a conversion

A PDF page contains text-showing operators: set this font at this size, move to this coordinate, draw this string of glyphs. That is the whole model. There is no paragraph object, no sentence, no reading order, no notion that two lines belong to the same block of prose. The information was discarded when the document was written, because a PDF only needs to know where to put ink.

A converter therefore reverses an inference. It groups glyph runs into lines by their vertical position, groups lines into paragraphs by their spacing and left edges, and guesses at headings from font size. Word then needs all of that, because a .docx describes paragraphs that reflow.

Where the guesses are right the output is genuinely editable. Where they are wrong you get the familiar symptoms: a paragraph broken at every visual line, a two-column page interleaved line by line, a header repeated as body text, a table rendered as rows of tab-separated text. None of these are defects in the reading of the PDF — the reading was accurate — they are the cost of having to invent structure.

02

Which documents convert well, and which do not

Single-column prose in one consistent font converts well: a letter, a memo, a report body, a policy document. The inference has an easy job because spacing is regular and there is one obvious reading order.

Documents converted from Word in the first place are the best case of all, because their layout came from the same paragraph-flow model you are trying to recover and the geometry still reflects it.

What converts badly: anything multi-column, anything heavily designed, forms, invoices and statements laid out as grids, and pages where text is positioned as labels around artwork. For these the honest assessment is that the output needs more repair than it saves, and reproducing the layout in Word from scratch while copying the text across is often faster.

03

When plain text extraction is the better tool

Plain text extraction takes the same glyph runs and writes them out as characters, without pretending to know about paragraphs. It has options for approximating the page layout with whitespace and for inserting page markers, and it makes no claim about structure.

Choose it whenever what you want is the words rather than the document: quoting a passage, feeding content into another system, searching a stack of files, or pulling figures out of a report. It never fabricates a structure that has to be unpicked later, and for those uses the formatting a Word conversion works so hard to produce is something you would have stripped anyway.

The decision comes down to a simple test. If you need to edit the document as a document — track changes, restyle it, send it for review — convert to Word and accept the repair work. If you need the content, extract text and skip the middle step.

04

Scanned PDFs need recognition first

If the pages are scans, there are no glyph runs at all. A page holding one large image has nothing that either conversion path can read, and the output is an empty or near-empty document. This is the most common reason a conversion appears to fail silently.

The fix is OCR: recognise the text, produce a searchable PDF with a real text layer, then convert that. Be aware you are now stacking two inferences, since OCR guesses the characters and the converter guesses the structure, so proofread the result against the scan rather than trusting it.

The selectable-text test at the start is worth repeating because it takes two seconds and it tells you which of these two workflows you are actually in.

While you follow this

MyPDFilles tools process files in this browser tab.

For a MyPDFilles tool, your document is read from your disk into this browser tab and the result is handed to your browser’s download mechanism. We do not make the same claim for external applications discussed in educational guides; check their own privacy and security information before using them.

Verify MyPDFilles requests in your Network panel

Questions about this task

Why is the formatting wrong in my Word file?

Because the PDF never stored formatting in the sense Word means — it stored positioned glyphs. The converter infers paragraphs, headings and emphasis from geometry and font names, and some of those inferences will be wrong on a complex page.

Do tables survive the conversion?

Often not as tables. Most PDF tables are drawn lines plus separately positioned text, with no table structure recorded. The text comes across in the right reading order but typically as lines rather than cells.

My converted document is empty. What happened?

Almost certainly a scanned PDF. Try selecting text in the original: if you cannot, the page is an image and there is nothing to extract. Run OCR first, then convert.

Should I choose Word or plain text?

Word if you need to edit and redistribute the document. Plain text if you need the content for quoting, searching or processing — it is faster, cleaner, and does not invent structure you would have to undo.