Extract text from a PDF
Get everything a PDF says as plain text you can search, paste or edit.
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Select a PDF file
or drag and drop here
Accepted formats: PDF
Processing…
Quick facts
- What it does
- Extracts the text layer from a PDF and returns it as a plain text file.
- Input
- Output
- TXT
- Processing
- In your browser
- Files uploaded?
- No
- Account required?
- No
- Best for
- Getting the words out of a report for quoting, searching or pasting elsewhere.
- Main limitation
- It reads the text already stored in the file. A scanned PDF that is only page images has no text layer, and this tool cannot create one — that needs OCR.
How to use this tool
- Upload your PDF.
- Choose whether to keep page breaks and page markers.
- Preview the extracted text.
- Download it as a .txt file, or copy it.
Why use iBuildPDF
Copying text out of a PDF by hand produces broken line endings, missing spaces and lost paragraphs. Reading the text layer properly gives you something you can actually work with.
How it works
The text layer of each page is read directly, with word and line positions used to rebuild spacing and line breaks. Nothing is re-typed or guessed — if the document has a text layer, this is what it says.
Frequently asked questions
Nothing came out — why?
Your PDF is almost certainly a scan: a picture of a page with no text layer at all. There is nothing to extract until the page has been through text recognition. That is what OCR does, and it is a different tool.
Does it keep the layout?
It keeps reading order, paragraphs and line breaks. It does not keep columns, tables or styling — plain text has no way to express those. For layout, PDF to Word is the right conversion.
Why did extracting text from my scan return nothing?
Because there is no text in it to extract. A scanned page or a photo saved as a PDF is an image, and extraction reads real text objects rather than recognising letters in a picture. Turning an image of words into actual text needs OCR, which is a different technology and one this site does not currently offer. Your file is not broken — it simply has no text layer.
Does the extracted text keep the original layout?
Not reliably, and this catches people out. A PDF positions text by coordinates rather than storing paragraphs, columns or tables as structures, so extraction has to infer reading order from position. Two-column layouts can interleave, and table cells can run together on one line. Extraction is the right tool for getting the words; it is the wrong tool for preserving how the page looked.
How can I tell whether my PDF has real text?
Open it in any PDF reader and try to select a few words with your cursor. If individual letters highlight and you can copy them, the text is real and extraction will work. If dragging selects the whole page as one block, or nothing highlights at all, it is an image and you are looking at a scan — no extraction tool will find text in it without OCR first.
Read more about this
Continue with your PDF
Your file is ready. These are the things people usually do next.