Skip to content

Convert From PDF

PDF to DOCX

A real Office Open XML .docx package built from your PDF’s extracted text.

Processed locally in your browser. Your file is not uploaded to any server. How this works

  1. 01Add your file
  2. 02PDF to DOCX
  3. 03Download

Runs in your browser · nothing uploaded

About PDF to DOCX

This page uses the same conversion engine as PDF to Word, with an additional control for inserting breaks between selected PDF pages. The output is a real .docx: a ZIP package containing word/document.xml for the body, word/styles.xml for paragraph styles, docProps/core.xml for document properties, and the relationship and content-type parts that connect them. It is not an HTML or RTF document with its extension changed. The package contains editable paragraphs, optional Heading 1–4 styles with outline levels, and optional explicit page breaks; it does not include a generated table of contents. The content is reconstructed from the PDF’s text layer. Text positions guide line grouping, large vertical gaps separate paragraphs, and adjacent body lines are joined, with some line-ending hyphens removed. Heading inference uses numbering, capitalisation and punctuation rather than source font size or weight. This conversion does not use PDF structure tags to recreate paragraphs, tables or a heading hierarchy. Source fonts, bold and italic, images and table grids are not transferred, and text from different columns can be combined. Explicit breaks separate selected-page content but do not guarantee the original pagination. Check the text, inferred structure and layout in your target editor or downstream reader; producing a real package is not a guarantee of identical behavior in every application or library. If the selected pages have no extractable text, conversion stops rather than creating an empty file.

How to PDF to DOCX

  1. 01

    Load the PDF

    Drop the file in. Parsing and packaging both happen locally.

  2. 02

    Set structural options

    Choose the pages, enable or disable text-based heading inference, and decide whether to insert breaks between selected-page sections. Breaks do not guarantee the original pagination.

  3. 03

    Generate the package

    The extracted lines are reconstructed into paragraphs and packaged as OOXML document, style, property and relationship parts in a ZIP archive.

  4. 04

    Download the .docx

    Save the .docx and check it in your target editor or OOXML reader. Review extracted text, inferred headings and page breaks before using it downstream.

What this tool does

  • A genuine Office Open XML ZIP package with document, style, property, relationship and content-type parts
  • Editable paragraph content inside the package, rather than a renamed HTML or RTF file
  • Optional explicit breaks between selected PDF pages, including pages with no extractable text
  • Inferred Heading 1–4 paragraph styles with outline levels for navigation and tables of contents you add in your editor
  • Generated in your browser, so confidential documents are never uploaded for conversion

Limitations worth knowing

Every PDF tool has constraints. Stating them plainly is more useful than discovering them halfway through your work.

  • The body is reconstructed from text positions, not copied from the original document structure. Reading order, spacing, paragraph joins and hyphen removal can need correction, especially on multi-column pages. Source fonts, bold, italic and page layout are not preserved.
  • Tables are not rebuilt as table elements or tab-separated grids. Cell text is grouped into ordinary lines and paragraphs, which can lose row and column relationships.
  • Embedded images are not carried into the package; this conversion works from the text layer.
  • If every selected page lacks extractable text, conversion reports a no-text-layer error and creates no document. For scanned text, use OCR PDF first. Pages with no extractable text in a mixed selection contribute no text and trigger a warning.
  • If you only need plain text, PDF to Text avoids the additional Word paragraph and heading reconstruction. It still uses text extraction, so check its reading order, spacing and content rather than assuming there is nothing to correct.

How your file is handled

This tool runs entirely inside this browser tab. When you choose a file, your browser reads it from your own disk and hands the bytes to JavaScript running on this page — no network request carries your document anywhere. You can confirm that yourself: open your browser’s developer tools, switch to the Network panel, and run the tool. You will see no upload.

Nothing is stored after the fact. Closing or reloading this tab discards the file, the result and everything derived from them, because none of it ever left your machine. Read how local processing works.

Questions about PDF to DOCX

How does this differ from the PDF to Word tool?

Both pages use the same engine and produce .docx files with the same text-reconstruction limits. This page exposes a working option to insert breaks between selected pages. The PDF to Word page instead shows a bold-and-italic option that currently has no effect.

Is the output a real .docx or a renamed file?

A real .docx. The ZIP contains OOXML document and style parts, document properties, relationships and content types. It is not renamed HTML or RTF. Check the generated package with your target editor or reader; this does not guarantee support for every application or library.

Will my original layout be preserved?

No. The tool reconstructs lines and paragraphs from text positions and can optionally infer headings from text patterns. Reading order and hierarchy can be wrong, particularly on multi-column pages. It does not reproduce table grids, images, source fonts, bold or italic, or the original page layout.

Should I turn on explicit page breaks?

Turn it on if you want an explicit break between the content extracted from each selected page. These breaks do not preserve source page numbers or guarantee the original pagination, because the text uses different styles and reflows. Leave it off for more continuous editing, and check any citations or cross-references against the PDF.

Can I get the images too?

Not from this tool. Run Extract PDF Images to pull the embedded raster assets out as files, then insert the ones you need.

Read more about this