How PDF compression actually works
What actually happens when a PDF is compressed?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
In almost every large PDF the bytes are photographs, so compressing the file means decoding those images, shrinking them, and writing them back. Text, vectors and structure are already compressed and are left alone unless you choose the destructive mode that rasterises whole pages.
The short answer
A PDF is not one compressed blob. It is a collection of numbered objects — page descriptions, fonts, images, annotations, metadata — held together by a cross-reference table, as defined in ISO 32000-1, and each object carries its own compression filter. There is no single dial marked “smaller”; there are several different things a tool can do, acting on completely different parts of the file.
Three of them matter in practice. Structural optimisation tidies the object graph and usually wins very little. Image recompression decodes the embedded photographs, reduces their pixel dimensions and re-encodes them, and this is where nearly all real savings come from. Rasterisation throws away the page description entirely and replaces each page with a picture of itself; it is the largest lever and the only destructive one. Knowing which of the three your file needs is most of the problem.
Where the bytes actually are
A 40-page text document is usually well under a megabyte. A 10-page scanned contract can be twenty times that. The difference is not the page count: one file stores instructions, the other stores photographs.
Where the bytes actually are
- Images. In almost every large PDF this is nearly the whole file.
- Embedded fonts — a few hundred kilobytes, and the same whether the document is 5 pages or 500.
- The text itself, and the structure holding it together. Usually a rounding error.
Text lives in a content stream: a short program of drawing operators saying which font, at which size, at which coordinates, with which characters. That stream is stored with the Flate filter — the DEFLATE algorithm of RFC 1951, the same compression used inside ZIP and PNG. It has already run. A page of prose costs a few kilobytes; forty of them cost a fraction of what one photograph costs.
A scanned page is the opposite. There is no text and no layout description at all: one image object, usually a JPEG, drawn across the whole page, and a content stream three or four operators long. The file is the pictures.
This is why “compress my PDF” is nearly always “recompress the pictures”, and why results vary so wildly between files that look similar. A report full of product photographs has enormous slack in it. A text-only PDF has almost none: run a general-purpose compressor over it and you will typically gain single-digit percentages, because the only thing left to squeeze has already been squeezed. If you do not know which kind of file you have, /pdf-page-count reports the page count, the page sizes and whether a text layer is present.
Lossless and structural optimisation
Before touching any image, a writer can reorganise the file itself. Nothing visible changes and nothing is lost. There are three common wins.
- Object streams. ISO 32000-1 allows many small non-stream objects to be packed into a single compressed stream rather than written one by one with their own overhead.
- Dropping orphaned objects. Editing a PDF often appends rather than rewrites. A page you deleted, a font you replaced, an earlier version of an annotation — these can still be sitting in the file, referenced by nothing. A full rewrite keeps only what the document actually points at.
- Stripping metadata. Document properties, XMP packets and application-specific private data can be removed. This is kilobytes, not megabytes, and its real value is privacy rather than size; /remove-pdf-metadata is the tool for that.
Be realistic about the payoff. On a PDF exported cleanly once from a modern application, structural optimisation typically wins single-digit percentages and sometimes almost nothing, because the exporter already did it. On a file opened, edited and re-saved a dozen times across several applications, orphaned objects can be a substantial share of the file and a rewrite wins considerably more. What it will never do is turn a 50 MB scan into a 2 MB one.
Recompressing the images: two independent levers
Once you accept that the images are the problem, there are exactly two things you can change about each one, and they are independent.
Resolution
An image’s effective resolution is not a property of the image alone. It is the pixel width divided by the width it is actually drawn at on the page. A 2400-pixel-wide photograph across an 8-inch page is 300 DPI. The same file placed in a one-inch corner as a logo is 2400 DPI, and every pixel beyond the first few hundred is invisible — detail no screen and no printer will ever show.
So the useful question is never “how big is this image” but “how big is it relative to how it is used”. Downsampling resizes the pixel grid so the effective resolution lands somewhere sensible. Around 150 DPI is comfortably above what a screen resolves and well short of what print needs; below about 100 DPI, text inside a scanned image visibly softens.
Quality
JPEG, specified in ITU-T T.81, is lossy by design: it transforms blocks of pixels into frequency coefficients and quantises them, discarding the detail the eye is least sensitive to. The quality setting controls how aggressively. Low quality is smaller and eventually produces block artefacts and haloing around sharp edges — which is exactly where scanned text lives, so scans tolerate quality reduction worse than photographs do.
What iBuildPDF does
Smart mode on /compress-pdf walks the PDF’s object graph, finds each embedded image XObject, decodes it, downsamples it to roughly 150 DPI based on how large it is actually drawn on the page, re-encodes it as JPEG at quality 0.80, and replaces that image stream in place. Strong mode is the same mechanism at roughly 96 DPI and quality 0.58.
Page content streams are never touched. The consequence matters: text stays selectable and searchable, vector artwork stays sharp at any zoom, and links, outlines, form fields and annotations all survive. Only the pixels inside image objects change. Some images are deliberately skipped — why some PDFs cannot be compressed covers which and why.
Rasterising: the destructive option
The last approach abandons the document model. Each page is rendered as a reader would see it, saved as a JPEG, and a new PDF is built containing nothing but those pictures.
It works where nothing else does, because it is indifferent to why the original was large. A dense vector map of a hundred thousand drawing operations becomes one flat photograph. A document whose images are all in formats that cannot be recompressed becomes a document with no such images. The output size depends only on the page dimensions and the JPEG quality, not on the original’s contents.
The cost is severe and permanent:
- No selectable text. Nothing can be copied out.
- No search, in any reader, ever.
- Links, bookmarks, form fields and annotations are gone.
- Screen readers have nothing to read; the document becomes inaccessible.
- Vector artwork is now fixed-resolution and blurs when zoomed or printed larger.
iBuildPDF’s Maximum mode does this, and it is the only destructive mode. The page labels it as destructive before you run it, so the choice is explicit rather than a surprise you find later. There is no way back: the text is not hidden, it no longer exists in the file. And because iBuildPDF does no OCR, no tool here can put a text layer back afterwards. Keep the original.
Choosing an approach
| Approach | What changes | Text stays selectable | Typical use |
|---|---|---|---|
| Structural / lossless | Object layout, orphaned objects, metadata. No visible change. | Yes | Files that have been edited and re-saved many times; privacy clean-up. |
| Image recompression (Smart, ~150 DPI, q0.80) | Embedded photographs are downsampled and re-encoded. Content streams untouched. | Yes | The default. Reports, brochures, anything with photographs in it. |
| Image recompression (Strong, ~96 DPI, q0.58) | Same mechanism, more aggressive. Visible softening on detailed images. | Yes | A hard size limit you cannot otherwise meet, where screen legibility is enough. |
| Rasterisation (Maximum) | Every page becomes a flat image. Everything else is discarded. | No | Last resort: vector-heavy files, or nothing else got close. |
A workable order: start with Smart and look at the result. If it is small enough, stop — you have lost nothing that matters. If it is not, work out why the file is large before escalating, because Strong mode only helps if the file is large because of photographs. For everything else, see why some PDFs cannot be compressed; if you are working to a fixed limit, getting a PDF under 1 MB covers what to do when the target is not reachable.
All of this runs in your browser: the file is read with the File API and processed locally by pdf-lib and PDF.js, and nothing is uploaded. The practical ceiling is your device’s memory rather than a fixed file size.
Tools this article covers
Sources
Primary documentation for the claims above.