Extract text from a PDF

Get everything a PDF says as plain text you can search, paste or edit.

Select a PDF file

or drag and drop here

Accepted formats: PDF

Private: processed on your deviceYour file never leaves your browser.

Quick facts

What it does
Extracts the text layer from a PDF and returns it as a plain text file.
Input
PDF
Output
TXT
Processing
In your browser
Files uploaded?
No
Account required?
No
Best for
Getting the words out of a report for quoting, searching or pasting elsewhere.
Main limitation
It reads the text already stored in the file. A scanned PDF that is only page images has no text layer, and this tool cannot create one — that needs OCR.

How to use this tool

  1. Upload your PDF.
  2. Choose whether to keep page breaks and page markers.
  3. Preview the extracted text.
  4. Download it as a .txt file, or copy it.

Why use iBuildPDF

Copying text out of a PDF by hand produces broken line endings, missing spaces and lost paragraphs. Reading the text layer properly gives you something you can actually work with.

How it works

The text layer of each page is read directly, with word and line positions used to rebuild spacing and line breaks. Nothing is re-typed or guessed — if the document has a text layer, this is what it says.

Frequently asked questions

Nothing came out — why?

Your PDF is almost certainly a scan: a picture of a page with no text layer at all. There is nothing to extract until the page has been through text recognition. That is what OCR does, and it is a different tool.

Does it keep the layout?

It keeps reading order, paragraphs and line breaks. It does not keep columns, tables or styling — plain text has no way to express those. For layout, PDF to Word is the right conversion.

Why did extracting text from my scan return nothing?

Because there is no text in it to extract. A scanned page or a photo saved as a PDF is an image, and extraction reads real text objects rather than recognising letters in a picture. Turning an image of words into actual text needs OCR, which is a different technology and one this site does not currently offer. Your file is not broken — it simply has no text layer.

Does the extracted text keep the original layout?

Not reliably, and this catches people out. A PDF positions text by coordinates rather than storing paragraphs, columns or tables as structures, so extraction has to infer reading order from position. Two-column layouts can interleave, and table cells can run together on one line. Extraction is the right tool for getting the words; it is the wrong tool for preserving how the page looked.

How can I tell whether my PDF has real text?

Open it in any PDF reader and try to select a few words with your cursor. If individual letters highlight and you can copy them, the text is real and extraction will work. If dragging selects the whole page as one block, or nothing highlights at all, it is an image and you are looking at a scan — no extraction tool will find text in it without OCR first.

Read more about this