PDF to text

Copy text from a PDF – without broken line breaks

A paragraph from a PDF needs to go into an e-mail, a Word document or a translation tool. After pasting, every line ends with a break, words are torn apart by “hyphen- ation” – or nothing can be selected at all. Here's how to get clean text.

Get started: open the PDF, copy the text as clean running text or save it as TXT – even from scans. Free, no upload.

Get text from PDF

Why copied PDF text is broken

A PDF is made for printing. It doesn't store “a paragraph with this text” but “these letters at exactly this position”. Lines, paragraphs and columns only follow from the layout. When copying, the app guesses where a line ends – and inserts a hard break there.

ProblemCause
A break after every lineEach line is a separate text piece in the PDF.
“implemen- tation”The layout's hyphenation is copied as a real hyphen.
Columns mixed upLines from the left and right column are read alternately.
Strange characters instead of textThe font is embedded without a character map – programs can't read the text.
Nothing can be selectedThe PDF is a scan (just an image) or copying is locked by a permission setting.

Clean text in three steps

  1. Open the PDF. Drop the file into PDF to Text. The text of all pages appears in the text box.
  2. Check the settings. “Remove line breaks inside paragraphs” and “Join hyphenated words” are on by default. Bullets and numbered lists are kept. Only some pages? Enter e.g. 3-5, 8 under “Pages”.
  3. Copy or save. “Copy text” puts everything on the clipboard, “Save as TXT” downloads a text file.

The PDF never leaves your computer. Contracts, reports or medical letters are safe to convert – the text is read in your browser.

Running text instead of line salad – drop the PDF in, copy, done.

When nothing can be selected

Scanned PDFs: text recognition

If the tool finds pages without text, a hint appears with the button “Recognize text (OCR)”. Recognition (Tesseract) also runs in your browser and reads German and English. The first time it downloads about 20 MB of recognition data; after that it's faster.

Alternatives: Word, Acrobat, pdftotext

Microsoft Word

Word opens PDFs directly and converts them into an editable document. Fine for simple layouts; complex pages end up with many text boxes and breaks.

Adobe Acrobat

“Save as text” and OCR are part of the paid versions. The free Reader only lets you select and copy.

Google Docs

Upload the PDF to Google Drive and open it with Google Docs – Drive runs OCR in the process. The file then lives at Google, though.

Command line

pdftotext from Poppler extracts text from many PDFs in seconds, ocrmypdf makes scans searchable.

pdftotext -layout document.pdf document.txt
ocrmypdf -l eng scan.pdf scan-searchable.pdf

Get the text out of your PDF

Clean running text without line salad, copy or save as TXT – even from scans via OCR. Free, no upload.

Frequently asked questions

Why does copied PDF text have so many line breaks?

PDFs store text line by line at fixed positions. When copying, every line end becomes a hard break.

How do I remove the line breaks?

With PDF to Text: the option “Remove line breaks inside paragraphs” joins lines back into paragraphs while keeping lists.

Can I copy text from a scanned PDF?

Yes, with OCR. The tool detects pages without text and reads them with OCR on request – right in your browser.

Is my PDF uploaded?

No. OCR runs locally too. Only the recognition data is downloaded once.

Is formatting kept?

No, the result is plain text without bold, font sizes or table lines.

What about multi-column layouts?

They're usually read in the right order, but complex layouts can mix up columns.

Can I convert only some pages?

Yes. Enter e.g. 1-3, 5 under “Pages”.

Sources

  1. Tesseract OCR and Tesseract.js – the text recognition the tool uses in your browser.
  2. Poppler (pdftotext) and OCRmyPDF.
  3. PDF.js – the open-source library the tool uses to read the text.

Copy text from PDFs without line salad – free, no upload.

Open the tool