File size and compression

Why some PDFs will not get any smaller

Why won’t my PDF get any smaller?

Short answer

Usually because there is nothing left to compress: the file is already efficiently built, or its size comes from embedded fonts, dense vector artwork, or image formats that cannot be recompressed in a browser. Image compression only helps a file whose bytes are photographs.

The short answer

Every compression tool has the same limit: it can only remove redundancy that is still there. When a PDF refuses to shrink, it is almost never that the tool failed. It is that the file has no slack in it, or that its bulk sits in a part of the document image compression does not reach.

Four causes are common, and they call for different responses. The file is already near its floor. Its size is embedded fonts. Its size is vector artwork. Or its images are in formats that cannot be decoded and re-encoded in a browser. Work out which you have first — /pdf-page-count reports the page count, the page sizes and whether a text layer is present, which is usually enough to tell.

The file is already efficiently built

This is the most common case and the least satisfying. A PDF exported once from a modern application is usually already close to as small as it can be, because the generic compression has already run.

123

Where the bytes actually are

  1. Images. In almost every large PDF this is nearly the whole file.
  2. Embedded fonts — a few hundred kilobytes, and the same whether the document is 5 pages or 500.
  3. The text itself, and the structure holding it together. Usually a rounding error.

Page content streams — the instructions describing where the text goes, which fonts to use, how to draw the lines — are stored with the Flate filter, the DEFLATE algorithm of RFC 1951, the same one inside ZIP and PNG. The exporter applied it when it wrote the file. Running DEFLATE over already-DEFLATEd data recovers nothing; compressed output looks like noise to a compressor, which is precisely why it is compressed.

So a 40-page text document with no images has no meaningful slack. Structural optimisation — repacking objects into object streams, dropping objects orphaned by earlier edits, stripping metadata — wins single-digit percentages and sometimes barely that. That is not a limitation of a particular tool; it is what the file is.

One exception is worth knowing. A PDF opened, edited and re-saved many times across several applications can be carrying a lot of unreferenced junk from earlier versions of itself, and rewriting it properly wins considerably more. If a file seems implausibly large for what it contains and has a long editing history, that is the likely explanation.

Embedded fonts

A PDF is supposed to look the same everywhere. That promise is kept by putting the fonts inside the file, so the document does not depend on what happens to be installed on the reader’s machine. It is one of the reasons the format exists, and it costs bytes.

How many depends on the font. A subset — only the glyphs the document actually uses — of a Latin text face is typically tens of kilobytes and not worth worrying about. A full, unsubsetted font is larger. A CJK font, covering Chinese, Japanese or Korean, can run to several megabytes on its own, because it holds many thousands of glyphs. Embed two weights of one and you have a multi-megabyte PDF containing a page of text.

Poor subsetting is a real and fixable inefficiency, but the fix is to re-export from the source document with subsetting enabled — a job for the application that made the PDF, not for a compressor.

What you should not do is strip the fonts out. A PDF whose fonts have been removed renders with substitutes on any machine lacking the originals: metrics shift, line breaks move, the layout drifts, and characters can be missing entirely. The document silently stops being the document. iBuildPDF does not strip embedded fonts, in any mode, for exactly this reason.

Dense vector artwork

A map, a CAD export, a floor plan, a chart with tens of thousands of plotted points: these look like pictures and are not. They are vector drawings, so the page content stream holds an individual drawing operation for every line, curve, fill and clipping path in the artwork. A detailed map can be hundreds of thousands of operations. That is the file.

There is no image quality lever to pull, because there is no image. Nothing to downsample, nothing to re-encode, no JPEG quality setting that applies. Run image compression over a file like this and it will correctly report that it found nothing to work on, and the output will be the same size as the input. That is not a failure; it is the honest result. The content stream is, as always, already Flate-compressed, so the generic lever has been pulled too.

The real options are to simplify the artwork in the application that produced it — fewer paths, less hidden detail, unneeded layers removed — or to give up the vectors and rasterise, which is covered below. The consolation is that vector files are resolution-independent: a 500 KB map prints perfectly at any size, which a 500 KB photograph of a map does not.

Image formats that cannot be recompressed

Sometimes the file genuinely is images and they still cannot be touched, because the browser cannot decode them. Recompression means decoding the original pixels first, and a browser-based tool has only the decoders the browser ships with.

  • JPEG 2000 (JPXDecode). A wavelet-based successor to JPEG, used in some archival, medical and professional scanning workflows. Browsers do not reliably support it — there is no dependable native JPX decoder to hand, so there is no safe way to get at the pixels.
  • JBIG2. A format for bilevel (pure black-and-white) scanned images. It recognises repeated shapes — the same letterform appearing a thousand times on a page — and stores each shape once. It is extremely efficient on scanned text and a JPEG could not improve on it.
  • CCITT Group 3 and Group 4 fax. The classic bilevel fax encodings, still produced by document scanners and multifunction printers. Each pixel is one bit, so the result is typically already smaller than any JPEG we could make from the same page; converting a crisp bilevel scan to JPEG would make it both larger and worse, with grey haloing around every letter.

Other things are skipped on purpose too: stencil masks and anything used as an /SMask or /Mask, images smaller than about 100×100 pixels, raw samples at anything other than 8 bits per component, colour spaces that cannot be mapped reliably such as Separation, DeviceN and Lab, and CMYK JPEGs on a browser that fails a start-up decode probe.

When iBuildPDF skips an image it counts the skip and reports it rather than quietly doing nothing. If your file barely changed and the result says most of its images were skipped, that is the explanation.

What actually helps

Match the remedy to the cause.

  • Already efficiently built. Accept it, or change what you are sending — split the document, or send a link rather than an attachment. No setting will help.
  • Heavily re-saved file. A clean rewrite drops accumulated orphaned objects; re-exporting from the source application does the same.
  • Fonts. Re-export with subsetting enabled. Do not remove the fonts.
  • Vectors. Simplify the drawing at source, or rasterise.
  • Skipped image formats. JBIG2 or CCITT means the file is likely already near its floor; leave it alone. For JPEG 2000, re-exporting from the scanning software with ordinary JPEG output usually gives a file that can then be compressed normally.

For a vector-heavy file, rasterising is the only remaining lever. Maximum mode on /compress-pdf renders each page and replaces it with a JPEG of itself, so the output size depends only on the page dimensions and the quality rather than on the artwork’s complexity. It works. It is also the only destructive mode, and the page says so before you run it: the result has no selectable text, no search, no copy-paste, no working links or form fields, nothing for a screen reader, and vector artwork that blurs when zoomed or printed larger. Since iBuildPDF does no OCR, nothing here can restore a text layer afterwards. Keep the original.

For the mechanics of what each mode changes, see how PDF compression works.

Tools this article covers

Sources

Primary documentation for the claims above.

Last reviewed: September 16, 2026

Published by iBuildPDF.

More from the knowledge base

Browse the knowledge base