Original research

PDF compression benchmark

How much smaller does a PDF actually get, and what does it cost? Six document types, measured against the shipped compression engine.

What this measured

Compression savings depend almost entirely on how much of a document is photographs. A text-only PDF lost 10.7% of its size; a brochure of oversized photographs lost 91.3%. In every case the extracted text was identical before and after.

What this does and does not establish

Read this before the numbers, because it determines what they are worth.

  • The corpus is synthetic. Every test document is generated by a seeded script, not taken from real work. That is a deliberate trade: real documents would be more representative and impossible for anyone else to verify, because nobody else can have the files. These can be rebuilt byte-for-byte by anyone.
  • Export settings were chosen to match real software, not to flatter the result — scans at about 200 DPI and JPEG quality 80, designed pages at quality 85. An earlier draft of this corpus used quality 95 with chroma subsampling disabled, which nothing in the world actually emits, and which made the reductions look roughly twice as good as they are.
  • These figures are not a prediction about your file. They establish the mechanism and the order of magnitude. A document whose images are already sized correctly for the page will save far less than one where a 12-megapixel photograph was dropped into a thumbnail frame.
  • One machine, one browser. Times are from headless Chromium on a Linux server. A phone will be slower, and how much slower depends on the device.

Method

Each document was loaded into a page running the same compress-engine.js that the live Compress PDF tool uses, at the same two non-destructive settings the tool offers. Each case was run three times and the median wall-clock time is reported, so a single garbage-collection pause is not the published figure. The destructive Maximum mode, which converts pages to images, is not measured here: its result is a picture of the document, and comparing that on file size alone would be misleading.

Run dateSeptember 16, 2026
Site version3.6.0
Engineassets/js/compress-engine.js
Librariespdf-lib 1.17.1
BrowserChromium (Playwright build), headless, Linux x86-64
Timingmedian of three runs, in-page wall clock
Settings measured compress.level.smart — images targeted at 150 DPI, JPEG quality 0.8
compress.level.strong — images targeted at 96 DPI, JPEG quality 0.58

The corpus

Six documents, chosen to span the cases that behave differently rather than to produce a flattering average. The checksum is the first 16 hex characters of the file's SHA-256, so a rebuild can be confirmed as identical.

Document Pages Size SHA-256
Text-heavy document
A 42-page agreement. Body text only, no images.
42 93 KB fa679b488de94d39
Scanned document
18 pages photographed at about 200 DPI and saved as JPEG at quality 80 — typical scanner and phone-scan output.
18 10.52 MB c9862d8ba23e0538
Photo-heavy brochure
12 pages, three photographs per page. Some placed at roughly 200 DPI, some at over 300 DPI for the frame they sit in.
12 4.12 MB 65f2832b02d32676
Presentation export
24 landscape slides, one photograph and a bullet list each.
24 2.87 MB b379395c565b93ae
Mixed text and images
30 report pages, one photograph per page at quality 85.
30 3.80 MB 5090d8357755733f
Already optimised
The output of the "mixed" document after one Smart pass, fed back in.
30 414 KB 1569fec27295704a

Results

"Images" is how many embedded images were recompressed out of how many the document contains. "Text" is the number of words extracted from the output with pdftotext, compared against the same extraction from the input.

Document Mode Before After Saved Time Images Text
Text-heavy document compress.level.smart 93 KB 83 KB 10.7% 0.03 s 0 / 0 25,074 words, unchanged
Text-heavy document compress.level.strong 93 KB 83 KB 10.7% 0.02 s 0 / 0 25,074 words, unchanged
Scanned document compress.level.smart 10.52 MB 4.17 MB 60.4% 1.41 s 18 / 18 No text layer
Scanned document compress.level.strong 10.52 MB 1.51 MB 85.7% 1.12 s 18 / 18 No text layer
Photo-heavy brochure compress.level.smart 4.12 MB 369 KB 91.3% 0.68 s 36 / 36 2,136 words, unchanged
Photo-heavy brochure compress.level.strong 4.12 MB 167 KB 96.0% 0.53 s 36 / 36 2,136 words, unchanged
Presentation export compress.level.smart 2.87 MB 329 KB 88.8% 0.55 s 24 / 24 912 words, unchanged
Presentation export compress.level.strong 2.87 MB 138 KB 95.3% 0.42 s 24 / 24 912 words, unchanged
Mixed text and images compress.level.smart 3.80 MB 414 KB 89.4% 0.67 s 30 / 30 9,990 words, unchanged
Mixed text and images compress.level.strong 3.80 MB 195 KB 95.0% 0.51 s 30 / 30 9,990 words, unchanged
Already optimised compress.level.smart 414 KB 414 KB 0.0% 0.21 s 0 / 30 9,990 words, unchanged
Already optimised compress.level.strong 414 KB 195 KB 52.8% 0.26 s 30 / 30 9,990 words, unchanged

Quality and integrity checks

Two checks were run on every output. The first renders page one of the input and the output at the same resolution and reports the mean absolute per-channel difference, on a 0–255 scale — a number near zero means the page looks the same, and a large number would mean something went wrong such as an inverted or blank image. The second runs qpdf --check to confirm the file is structurally valid rather than merely openable.

Document Mode Page 1 pixel difference Structure
Text-heavy document compress.level.smart 0.00 / 255 Valid
Text-heavy document compress.level.strong 0.00 / 255 Valid
Scanned document compress.level.smart 2.50 / 255 Valid
Scanned document compress.level.strong 4.93 / 255 Valid
Photo-heavy brochure compress.level.smart 0.47 / 255 Valid
Photo-heavy brochure compress.level.strong 0.81 / 255 Valid
Presentation export compress.level.smart 0.35 / 255 Valid
Presentation export compress.level.strong 0.78 / 255 Valid
Mixed text and images compress.level.smart 0.27 / 255 Valid
Mixed text and images compress.level.strong 0.42 / 255 Valid
Already optimised compress.level.smart 0.00 / 255 Valid
Already optimised compress.level.strong 0.40 / 255 Valid

Observations

The photographs are the file. The text-only agreement is 42 pages and under 100 KB; the 18-page scan is 10.5 MB. Page count is almost irrelevant to PDF file size. What matters is whether the pages are instructions or pictures.

Over-resolution is where the savings come from. The brochure lost 91.3% not because the images were badly compressed but because they were far larger than the space they were drawn in — a 1100-pixel-wide photograph placed in a 230-point frame is about 344 DPI, and reducing it to 150 DPI removes roughly four fifths of its pixels before any quality change. A document whose images were already sized for the page has much less to give.

Text survived every non-destructive run. All six documents that had a text layer extracted exactly the same word count afterwards, which is the expected result of replacing image streams while leaving page content streams untouched — but it is worth measuring rather than assuming.

Compressing twice gains nothing. Feeding a Smart-compressed file back through Smart produced no change at all: all 30 images were examined and none could be made smaller, so the original was kept. Running the same file at the Strong setting did reduce it further, because that is a genuinely lower resolution and quality target, not a second pass at the same one.

A scan is the hardest case. It has no text to protect and nothing but images to work on, so it responds well in percentage terms — but it starts far larger, so the result is still the biggest file in the set.

Reproducing this

The corpus generator, the measurement harness and the verification script ship with the site under tools/: benchmark-corpus.py builds the six documents from fixed seeds, benchmark-run.mjs drives the engine in a real browser, and benchmark-verify.py runs the text-extraction and rendering checks. The generated results in config/benchmarks.php are what this page reads — no figure on this page is written by hand, and changing the engine without re-running the harness leaves the old run's date visible rather than producing new claims.

Tools this article covers

Read more about this