Is this PDF scanned or is it text?
How can I tell whether a PDF is scanned?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Try to select a word with your cursor. If the text highlights, the file contains real text. If you get a selection box over a picture instead, the page is an image and the file contains no text at all. The two look identical on screen and behave completely differently in every tool you point at them.
The two-second test
Open the file and try to drag-select a word.
- The word highlights — there is a text layer. Search, copy and conversion will all work.
- A rectangle drags across the page instead, selecting nothing — the page is one image. There is no text to find.
A second check: press Ctrl+F (Cmd+F) and search for a word you can plainly see. No match on a word that is visibly there means no text layer.
The same page, with and without text
- A digital PDF. On screen, indistinguishable from the one on the right.
- Underneath it: character codes, each with a font and a position. Searchable, selectable, convertible.
- A scanned PDF. The same page, and the same appearance.
- Underneath it: a grid of light and dark samples. No letters, no words, nothing to search.
Both pages above are the same page, and on screen they are indistinguishable. Underneath, one is a sequence of character codes with positions and a font; the other is a grid of light and dark samples that happens to look like writing to a human eye. Nothing in the second one knows it contains letters.
If you would rather not squint at it: run the file through PDF to text. A scanned file comes back empty or nearly so, and that is your answer.
Why it changes everything downstream
Almost every complaint about PDF tools traces back to this distinction.
- Convert to Word and you get photographs. Not a bug, and not something a better converter would fix: there was no text in the file to convert. The converter faithfully moved across what was there, which was two pictures.
- Search finds nothing, in the file and in any system that indexes it.
- Copy and paste gives you nothing.
- Compression behaves differently. A scan is image data, so it responds well to image compression and not at all to the tricks that shrink text — which is why some files refuse to get smaller and scans usually do.
- Redaction is a different job. There is no text to remove, only pixels to paint over — but the original image is still underneath unless the tool rebuilds the page.
- Screen readers get nothing at all. A scanned document is, to assistive technology, a blank page.
What OCR does about it
Optical character recognition looks at the picture, works out which shapes are letters, and writes a text layer underneath the image. The page still looks exactly the same — you are still seeing the scan — but now there is searchable, selectable, copyable text sitting invisibly behind it.
That is the operation OCR PDF performs, and it is the step that has to come first. Run OCR, then convert, search or extract; the tools that were useless on the original will work on the result.
Choosing the right language matters more than people expect. OCR decides between similar shapes using a model of what words look like in a particular language, so telling it the wrong one measurably increases errors — and for Arabic, Chinese or Greek, the wrong setting produces nonsense rather than mistakes.
What OCR will not fix
OCR is recognition, and recognition can be wrong. It is very good on a clean, straight, 300 DPI scan of ordinary printed text. It degrades on everything else, and it does not tell you when it is guessing.
- Handwriting is mostly beyond it.
- Low-resolution scans — below about 200 DPI — lose the detail that distinguishes similar letters.
- Skewed, shadowed or photographed pages do far worse than flatbed scans.
- Tables and multi-column layouts may be read in the wrong order, because the reading order has to be inferred from position.
- Stamps, signatures and handwritten annotations over printed text confuse the letter shapes underneath them.
So treat an OCR result as a searchable copy, not as a verified transcription. For anything where a wrong digit matters — an invoice total, a dosage, a legal reference — read the recognised text against the image before relying on it. OCR or PDF to Word? covers which of the two you actually want.
Tools this article covers
Sources
Primary documentation for the claims above.