Converting and extracting

How to get text and images out of a PDF

What is the best way to get content out of a PDF?

Short answer

It depends on what you want to do with it afterwards. For the words, extract the text. For the pages as pictures, render them to images. For a scan, recognise the text first — there is nothing to extract until you do. For something you will keep editing, convert to Word. Each gives a genuinely different thing, and picking by what you need next saves undoing the wrong one.

Getting the words

Extract text from a PDF reads the text layer and gives you plain text. No formatting, no images, no layout — just the words, which is exactly right when you want to quote, search, translate or paste into something else.

What you get back is shaped by how the PDF was made, and two things are worth expecting:

  • The reading order is the order the text was drawn, which for a multi-column page may interleave the columns. This is a property of the file rather than a fault in the extraction — see what makes a PDF accessible, where the same issue decides whether a screen reader can read a document.
  • An empty result means there is no text. Not that extraction failed — that the page is an image. This is the fastest way to find out whether a document is a scan.

Getting pictures of the pages

PDF to JPG and PDF to PNG render each page as it appears and give you one image per page. That is what you want for a thumbnail, a preview, a slide, or anywhere an image is easier to place than a PDF.

Choose the format by what is on the page: JPG for scans and photographs, PNG for text, diagrams and line art, where JPG artefacts land directly on the letterforms. JPG or PNG for a PDF goes into why.

Be clear about the trade first: rendering a page destroys its text. The result is pixels — not searchable, not selectable, not readable by a screen reader, usually larger, and not reversible. Right for a picture of a page; wrong for a document someone needs to read or reuse.

A related question this does not answer: these tools render whole pages, not the individual photographs embedded inside them. If you need the original images as they were stored, that is a different operation and a desktop tool is usually the answer.

Choosing, and the scan case

If the PDF is a scan, start with OCR. A scanned page has no text in it, so extracting text returns nothing and converting to Word returns photographs. OCR recognises the characters and adds a real text layer; everything else becomes possible afterwards. Recognition is never perfect and needs proofreading, and it recovers the words rather than the layout.

1234Tx

Four ways out, and what each gives you

  1. Plain text. Fast and exact, and it keeps none of the layout.
  2. Images of the pages. Everything is preserved and nothing is selectable.
  3. OCR, for a scan. Recognised text, with the recognition errors that implies.
  4. An editable document. The most useful and the least faithful, because the structure has to be inferred.

Otherwise, pick by the destination:

Text extraction and page rendering both run in your browser. Conversions to Word and Excel run on our server: the file is sent over an encrypted connection, held only while it is converted, and deleted afterwards.

Tools this article covers

Sources

Primary documentation for the claims above.

Last reviewed: September 24, 2026

Published by iBuildPDF.

More from the knowledge base

Browse the knowledge base