Why a scanned PDF is so much larger than the original
Why is my scanned PDF so large?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Because a scan is not a document, it is a set of photographs — one per page, each of several million pixels — and the page count matters far less than the resolution and colour mode they were captured at. Doubling the scanning resolution quadruples the data, and scanning in colour rather than grey triples it again before compression, which in a finished file usually lands as roughly two to three times the size. That is how ten pages of plain text become a file larger than a film.
Do the arithmetic once and it stops being mysterious
A scanner does not read a document. It samples a sheet of paper on a grid and stores the result, and the size of that grid is set by two numbers you chose without thinking about them.
An A4 page is 210 × 297 mm, which is about 8.3 × 11.7 inches. At 300 samples per inch that is roughly 2,480 × 3,500 pixels — about 8.7 million of them, for one page. In full colour, at three bytes a pixel, that is around 26 MB of raw data before any compression at all. Raise the setting to 600 DPI and the grid doubles in each direction, so the pixel count quadruples: about 35 million pixels and 104 MB raw, for the same sheet of paper.
Compression cuts those numbers down enormously — that is why a scanned page is typically a few hundred kilobytes rather than 26 MB — but it does not change the ratios. A 600 DPI scan is roughly four times the work of a 300 DPI scan whatever the compressor does afterwards, and a colour scan carries roughly three times the samples of a grey one.
Set that against the alternative. The same page written as text is a content stream of drawing instructions: which font, which size, which characters, at which coordinates. A page of prose costs a few kilobytes. That is the whole gap, and it is why a 40-page report exported from a word processor is smaller than a 4-page scan of the same report.
To find out which kind of file you actually have, page count and file info reports the page sizes and whether there is a text layer, and is my PDF scanned or text covers the two-second version of the test.
The two settings that decide everything
Resolution
Resolution only means something relative to how the result will be used. Printed text at 300 DPI is indistinguishable from the paper at arm’s length; on a screen, where you are looking at perhaps 100 to 150 pixels per inch of displayed page, half of that is already invisible. 600 DPI is for reproducing fine detail — a photograph you intend to print again, an engraving, handwriting you need to examine — and it is the wrong setting for an office document, where it quadruples the file in exchange for detail no reader will ever see.
What resolution actually decides
- 72 DPI. Fine on a screen, visibly stepped on paper.
- 150 DPI. Acceptable for an office printer and a document nobody will study.
- 300 DPI. What a commercial printer expects, measured at the size the image is actually printed.
Colour mode
Every pixel is stored at some depth. Full colour keeps three channels; greyscale keeps one; black-and-white keeps a single bit, on or off.
What the colour setting costs, per pixel
- Colour — three channels for every pixel. Worth it only when the colour carries information.
- Grey — one value per pixel, 256 levels. Enough for printed text, pencil, stamps and faces.
- Black and white — one bit per pixel, on or off. Very small, and everything in between is thrown away before it reaches the file.
After compression the ratios are not exactly 24 : 8 : 1 — JPEG treats colour information more coarsely than brightness, so a colour scan is often closer to twice a grey one than three times it, and a black-and-white scan compressed with the fax and bi-level schemes PDF supports can be smaller than a twentieth. The direction is reliable even where the exact factor is not.
| Setting | Good for | What it costs you |
|---|---|---|
| 150 DPI, grey | Reading on a screen, sharing a reference copy. | Small type softens. Not good enough to print from. |
| 300 DPI, grey | The sensible default for an office document: text, forms, signatures, pencil notes, stamps. | Nothing, for a document that was black ink on white paper. |
| 300 DPI, colour | Anything where the colour carries meaning — a highlighted contract, a coloured chart, a photograph. | Roughly two to three times the size of the same scan in grey. |
| 600 DPI, colour | Reproduction: photographs, artwork, anything that will be printed again at size. | Four times the pixels of 300 DPI. Almost never the right choice for a text document. |
| Black and white (1-bit) | Clean printed text and line drawings, and nothing else. | Everything between black and white is discarded. See below. |
Scanner menus rarely use these words. “Document”, “Text” or “Fax” modes are usually greyscale or 1-bit; “Photo” is colour at a high resolution. If yours offers a quality slider and nothing else, the middle of it is almost always the right answer for paperwork.
Fixing the file you already have
A scan is the one kind of PDF where compression reliably works, because a scan is nothing but images and images are where the bytes are. The order that gets the most out of it:
- Look first. Page count and file info, then divide the file size by the page count. Over about a megabyte a page and you are looking at a high-resolution colour scan with plenty of slack in it. Under about 150 KB a page and there is very little left to take.
- Drop the colour if it carries nothing. PDF to grayscale converts the pages to grey, which removes the colour channels and also gives the compressor an easier job afterwards. Do this before compressing, not after.
- Compress. Compress PDF decodes each embedded image, downsamples it to a sensible resolution for the size it is drawn at, and re-encodes it. How PDF compression works describes what each mode changes. If you are working to a fixed number, compress to 1 MB targets it directly.
Two honest warnings about compressing a scan specifically.
Scans tolerate it worse than photographs do. JPEG discards detail where the eye is least sensitive, and concentrates its errors around sharp high-contrast edges — which is exactly what a letterform is. A holiday photograph at low quality looks slightly soft; a scan of a page at the same setting grows fuzz around the letters and starts to look like a bad photocopy. Check the result at 100% before you send it, and why is my PDF blurry covers what you are looking at if it has gone wrong.
You cannot recover what the scanner discarded. Compression can throw away the excess in a 600 DPI colour scan and lose nothing you would notice at reading size. It cannot add detail back to a scan that was made at 150 DPI in the first place, and no setting on this site or anywhere else will.
When rescanning is simply the right answer
If you still have the paper, rescanning at the right setting is usually faster than fighting the file, and the result is better — because the compression artefacts were never introduced rather than being introduced and then partly removed.
For ordinary paperwork: 300 DPI, greyscale, scanned straight to a single multi-page PDF. That is small, sharp, prints correctly, keeps pencil and signatures legible, and needs no further work. Colour only when the colour means something.
Scanning each page to its own file and joining them afterwards with merge PDF works perfectly well and is sometimes the only option your scanner offers, but a scanner that can produce one document directly will usually produce a smaller one, because it compresses the whole job with consistent settings.
Phone photographs of pages are the case to avoid if you have any choice. Uneven lighting, a slight angle and a background all give the compressor more variation to preserve than a flatbed does, so the file comes out larger and harder to read. If a phone is all you have, a scanning app that flattens the perspective and thresholds the page is markedly better than the camera roll.
Three things worth knowing before you decide
Size is the least of what a scan costs you. A scanned page contains no text, so the document cannot be searched, cannot be copied from, cannot be read aloud by a screen reader, and converts to Word as a picture. OCR recognises the characters and adds a text layer underneath the image, which fixes all of that — and makes the file slightly larger rather than smaller, since the text is added to the picture rather than replacing it. That tool runs on a server, so the document is uploaded to be processed; for a confidential scan that is a decision to make deliberately.
The most aggressive black-and-white modes can change the characters. Some scanners compress 1-bit output by finding repeated shapes and storing one copy of each, which is extremely efficient and, where two characters look alike at the scanned resolution, has been documented to substitute one for the other — a digit quietly becoming a different digit in a table nobody re-read. If the numbers matter, do not use the smallest black-and-white setting, and compare the output against the paper.
What black-and-white scanning throws away
- A faint mark — a pencil note, a light stamp, a grey rule. On the paper it is plainly there.
- The scanner picks a threshold. Anything darker than the line becomes black, anything lighter becomes white, and there is nothing in between to store.
- The printed text survives because it was solidly dark. The faint mark is gone, and nothing done to the file afterwards brings it back.
Keep the original. Every step here is one-way. Downsampling, greyscale conversion and recompression all discard information permanently, and the moment to find out you needed the colour version is never the moment you expect.
Tools this article covers
Sources
Primary documentation for the claims above.