Skip to content

PDF to Word

Turn a PDF into an editable Word document.

or drop files anywhere on this page

Your files stay in this browser. Nothing is uploaded.

About PDF to Word

This tool converts a PDF to Word without uploading it. The PDF is opened and the .docx is built by code running in your browser tab, and neither file is sent to a server. What you get is text: paragraphs, headings, tables and links rebuilt as real Word content, not a picture of each page.

Reach for it when you need to quote, rewrite or repurpose what a PDF says. It guesses structure from where lines sit on the page, so a text-heavy contract converts well and a magazine spread with overlapping columns does not. Images are left out, and a scanned page has no text to convert until it has been through OCR, which this tool can run first.

How it works

  1. Add one or more PDFs. If a PDF needs a password to open, a password box appears beside the file and Convert to Word stays disabled until the password is accepted.
  2. Under Layout, choose Keep layout for one Word paragraph per line of the PDF with a page break after each page, or Optimize for legibility to merge lines into flowing paragraphs. Choose legibility to edit the text.
  3. If the PDF is a scan, the tool says it looks like one and shows a Run OCR first checkbox. Tick it and pick the document’s language under OCR language.
  4. Click Convert to Word.
  5. Download the .docx, which is named after the PDF, or a ZIP when you added more than one PDF. Check the headings and tables first.

What comes across, and what does not

The converter reads the text the PDF holds, one line at a time, and rebuilds structure from each line’s position, size and font name. It never looks at the page as a picture, so anything that is not text is left behind.

  • Headings: a line at least 1.8 times the document’s usual text size becomes Heading 1, and at least 1.25 times becomes Heading 2. There are no smaller levels.
  • Tables: two or more consecutive lines that break into separate columns become a Word table with equal-width columns. Cell text keeps no bold, italic or link, and two body lines that each change font midway, from regular into bold, can be read as a two-row table.
  • Links: clickable web addresses become hyperlinks in paragraphs and headings. Links to another page of the PDF are not kept.
  • Bold and italic are read from the font’s name, so a typeface not named Bold or Italic loses them.
  • Left out: images, drawings, colour, font sizes other than headings, alignment and indents. Running headers, footers and page numbers are not recognised as such and repeat as ordinary lines. Text typed into fillable form fields does not come across.

Choosing between Keep layout and Optimize for legibility

Keep layout writes each PDF line as its own Word paragraph and ends every page with a page break. Nothing reflows: type at the end of a line and the text will not wrap onto the next. The last page gets a break too, so you may find one empty page at the end.

Optimize for legibility joins the lines between headings and tables into one paragraph, with no page breaks, so text reflows as you edit. It cannot see where a paragraph ends, so a page’s consecutive body lines become one paragraph even when the PDF had several. A line-end hyphen followed by a lowercase letter is dropped and the halves joined, which also fuses a genuine hyphenated word split across two lines.

Both modes use A4 pages with 1 inch margins, whatever the PDF’s page size.

Scanned PDFs and the OCR option

A scan is a picture of a page, so there is no text to extract. When none of the first three pages has selectable text, the tool notes that the PDF looks like a scan. Convert without ticking Run OCR first and the Word file has nothing in it. Only those three pages are checked, so a file with text up front and scans behind does not offer OCR.

With Run OCR first ticked, every page is drawn at 200 dpi and read by an OCR engine in your browser, then converted like any other text. A language’s data is downloaded from a public CDN the first time you use it; your PDF is not part of that request. For scripts such as Cyrillic, Arabic, Hindi or Chinese, a font file is also requested from Google Fonts by name, and any word that font cannot draw is left out.

Proofread for misread characters. OCR text carries no bold or italic, because it is laid down in one plain font.

Frequently asked questions

Related tools

This tool ran entirely in your browser and nothing was uploaded. Your Recent files list keeps a copy of local files under 5 MB until you clear it. How your files are handled.