Skip to content
FreeConvertter

How to extract text from a PDF — and how to spot a scanned PDF

Published:

You copy a paragraph from a PDF and every line arrives as a separate fragment — or the text can’t be selected at all. Once you know how PDFs store text, the fix is straightforward.

There are two kinds of PDF

1. Text PDFs

Created by “Save as PDF” or “Print to PDF” from Word, Google Docs, Excel or a browser. The characters are stored as real text (a text layer), so you can select, search and copy them.

2. Scanned (image) PDFs

Made by scanning paper or photographing it with a phone. Each page is a picture; the letters look like text but are just pixels to the computer. You can’t select or search them.

How to tell: press Ctrl+F (or Cmd+F) in your PDF viewer and search for a word you can see on the page. If it’s found, it’s a text PDF; if not, it’s probably scanned. Being unable to drag-select any text is another giveaway.

Extracting text cleanly from a text PDF

Drop the file into the PDF to TXT converter and click Convert — the text of every page is saved as a plain-text file.

The converter rebuilds lines from the positions of the text fragments in the PDF:

  • fragments on the same baseline become one line,
  • a normal line spacing starts a new line,
  • a larger gap, as between paragraphs, inserts a blank line.

Under Options → Between pages, choose the page-number marker to insert lines like --- 2 ---, making it easy to find where a passage came from.

PDFs in Korean, Japanese and Chinese are supported too: the character maps (CMaps) these PDFs rely on are included.

When the result isn’t what you expected

Symptom Why, and what to do
“No text layer was found” It’s a scanned PDF — see below
Two-column papers come out interleaved PDFs don’t record columns, only positions. Copying column by column can be more accurate
Tables come out as runs of words PDF tables are drawn lines plus text; there are no cells to recover
Text turns into odd symbols The PDF uses fonts without proper character mappings. Extract from the original document if you have it

Have the original? Extract from that

A PDF records what printed pages look like; paragraph and table structure is gone. If you have the Word file the PDF came from, DOCX to TXT gives you far more accurate paragraphs and tables.

What about scanned PDFs?

Getting text out of a scanned PDF requires OCR (optical character recognition), which this site doesn’t offer yet. Meanwhile:

  • Many phones can recognise and copy text in photos. Turn the pages into images with PDF to PNG and use that feature.
  • Rescan the document with your scanner’s “searchable PDF” option.

Summary

If you can’t extract text from a PDF, first check whether it’s scanned. For text PDFs, PDF to TXT extracts everything with clean line and paragraph breaks — and when the original document is available, extracting from that is the most accurate option.

Tools used in this guide

Related guides