Why a PDF gets bigger when you save or edit it
Why did my PDF get bigger after I edited it?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Usually because the save appended rather than replaced. A PDF can be updated by leaving the original bytes alone and writing a new revision on the end, so the old copy of everything you changed is still in the file and still paid for. The other common cause is an edit that redrew the page as a picture — a flatten that could not use the form fields, a greyscale conversion, a redaction — which swaps a few kilobytes of drawing instructions for a few hundred kilobytes of pixels.
A save that adds rather than replaces
This is the cause people never guess, because nothing about it is visible. A PDF may be written as an incremental update: the file you started with is left exactly as it was, and a new section is appended to the end containing only the objects that changed, followed by a new cross-reference table that points back at the old one.
A save that appends, and the same file rewritten
- The file as first written: its objects, then a table saying where each of them is.
- A change saved by appending. The original bytes are untouched and the new revision holds only what changed — so the old copy of the changed object is still in the file, and still costs what it costs.
- The new table points back at the old one, so a reader knows where to look for everything this revision did not replace. Save five times and there are five tables in the chain.
- Rewriting the document instead of appending to it drops whatever the later revisions superseded. That is how a file comes back from a tool smaller than it went in without anything having been compressed.
ISO 32000-1 provides for this, and applications do it because it is fast — a one-word change to a four-hundred-page document is a few hundred bytes written to the end of a file rather than four hundred pages rewritten. It is also how a digital signature survives an annotation being added afterwards, since the signed bytes are still there, untouched, at the front.
The cost is that the old version of every object you changed is still in the file. Save five times and a document can carry five revisions, each holding the superseded copy of whatever the next one replaced. The document that opens is correct; the file is the history of it.
Two consequences, one about size and one that is more serious.
The size: a file can double over an afternoon’s editing without a single new page. Nothing is wrong with it and no compressor will help, because the bytes are not badly compressed — they are simply not needed.
The other: anything you removed and then saved incrementally may still be recoverable from the file. If you deleted a paragraph, a page or an image because it should not be shared, removing it from the current revision does not remove it from the bytes. How to redact a PDF permanently is the article about that, and it matters more than the megabytes.
Worth knowing about the tools here: they do not append. Nearly every tool on this site builds a new document and copies the pages into it, which means the output has one revision and whatever the old ones superseded is dropped. That is not a feature anyone designed for size — it exists because the PDF library this site uses writes a broken cross-reference table when it re-saves an incrementally-updated source, and Acrobat refuses the result. Measured on one real document while fixing that: 217,913 bytes through the re-save path against 192,994 bytes through the rebuild, with all 5,865 characters of text still extractable. A file that comes back smaller from a tool that did no compressing has usually just lost its revision history.
An edit that turned the page into a picture
The second common cause, and the one with a permanent cost attached.
A page of text is a list of instructions and costs a few kilobytes. The same page rendered and stored as an image is several million pixels, and no amount of compression brings it back to a few kilobytes. Any operation that rasterises a page therefore makes the file bigger — often much bigger — and takes the text layer with it.
Why rasterising a page makes the file bigger
- A page of text: a list of instructions — this font, this size, these characters, at these coordinates. A page of prose costs a few kilobytes.
- The same page rendered and stored as an image. Every letterform is pixels now, and the pixels are what the file has to carry from here on.
- The text goes with it: nothing to search, select, copy or read aloud. It is the one change that makes a document bigger and less useful at the same time.
Which operations on this site do that, exactly:
- Flatten PDF — only sometimes. It first tries to flatten the form fields directly, which keeps the text live and the file small. If that fails — appearance streams it cannot generate, annotations outside the form, an XFA form — it renders every page instead and rebuilds the document from JPEGs. The result page tells you which route ran, and that is not a detail: “your form is locked” means something different depending on whether the document is still searchable.
- PDF to grayscale — always. Every page is rendered, desaturated pixel by pixel and written back as a JPEG page. The three quality settings render at 2.0, 1.5 and 1.1 times the page size and encode at JPEG quality 0.92, 0.82 and 0.68. Put a text-only document through it and you should expect a larger file with no selectable text; the tool page says so before the upload box. It is the right tool for a colour scan and the wrong one for a report.
- Redact PDF — only the pages you marked. Marked pages are rendered at twice their size with the black boxes painted into the pixels, which is what makes a redaction permanent. Pages you did not touch are copied across with their text intact, so redaction costs you only the pages you redacted.
- Compress PDF at its most aggressive setting. The two lower levels recompress the images inside the document and leave it a document. The Maximum level rasterises, and it is labelled destructive on the page before you run it.
One limit that surprises people: a browser canvas cannot be arbitrarily large, so the render is capped at 4,096 pixels on a side and about 16 megapixels in total. A very large page is rendered at less than the requested scale rather than coming back blank, which means an A0 poster put through a rasterising tool comes out softer than an A4 page would. How large a PDF a browser can handle covers the other limits of working in a tab.
Fonts, images and the things that get embedded twice
The remaining causes are all the same shape: something that was shared, or partial, or already compressed, became unshared, complete or recompressed.
- A font that was subset gets embedded in full. A producer that embeds only the glyphs a document uses writes a few kilobytes; an editor that re-embeds the whole typeface writes a few hundred. The page looks identical. Why PDFs look different on other computers explains subsetting and its other surprises.
- The same font ends up in the file more than once. Merging is the usual way: each document’s pages are copied in along with the resources they reference, and nothing in the merge looks for two identical fonts and keeps one. Ten reports from the same template, all embedding the same typeface, carry ten copies of it into the merged file. How to combine PDF files covers what else changes.
- An image gets re-encoded losslessly. A JPEG photograph inside a PDF is already compressed. An editor that decodes it, changes something and writes it back as a lossless image can multiply its size while making it look very slightly better, which nobody asked for.
- A page is re-rendered by a print driver. Printing a PDF back to PDF is a complete re-render, and it frequently grows the file as well as stripping the document’s structure. What “Print to PDF” actually does is about that in full.
And one addition of a few bytes that is worth naming because it is this site’s own doing. iBuildPDF writes classic cross-reference tables rather than the compressed cross-reference streams the library would use by default. Both are valid; Acrobat frequently refuses the second, with “The root object is missing or invalid”, and a file the most common reader will not open is not a smaller file. Measured, the classic table costs about 2.7% on a realistic image-heavy document and between 17% and 27% on a tiny text-only one, where the table is most of the file. That is a deliberate trade and it is the whole of this site’s contribution to your file being larger.
Adding page numbers or a watermark, by contrast, costs almost nothing: both draw with one of the fourteen standard fonts, so no font data is embedded at all — a few hundred bytes per page of drawing instructions and no more.
Getting the size back
- Find out what kind of file you have. Page count and file info, then divide the size by the page count. A text document over a few hundred kilobytes a page has something in it that is not text, and knowing whether there is still a text layer tells you immediately whether a page has already been rasterised.
- If the text layer is gone, stop and go back. Rasterising is one-way. Compression will shrink the images happily, and the text is not coming back except through OCR, which is a fresh set of guesses rather than the original characters and runs on a server. The right move is the earlier version of the document, if you still have it.
- If the text layer is intact, compress. Compress PDF decodes the embedded images, downsamples them to a sensible resolution for the size they are drawn at and re-encodes them. It will not hand you a larger file: if its output would be bigger than what you gave it, it says the document was already optimised and gives you your own file back unchanged. How PDF compression works describes what each level changes.
- If the bloat is revision history, any structural tool will drop it. There is no “rewrite this as one revision” button here, and there is no need for one: rotating, reordering, extracting or compressing all write a fresh single-revision document, so the saving arrives alongside whatever else you were doing. Organize pages is the least invasive of them.
- Keep the file you started with. Every step above discards something permanently, and the moment you discover you needed the original is never the moment you expect.
If the file is over about a megabyte a page and is genuinely scanned, the arithmetic and the fixes are different enough to have their own article: why a scanned PDF is so large. And if nothing shrinks it at all, why some PDFs cannot be compressed covers the files that are already as small as they get.
Four things worth knowing before you try
- A bigger file is not a broken file. A document carrying four revisions opens correctly, prints correctly and is perfectly valid PDF. If nothing depends on the size, nothing needs doing.
- Rewriting a signed document breaks the signature. A digital signature covers a range of bytes, so any tool that writes a new file invalidates it — including every tool here, since they all rebuild. Incremental updates exist partly so that signed documents can be annotated without that happening. If the signature is the point of the document, do not put it through anything. How to sign a PDF covers the difference between a signature and a picture of one.
- Some growth is simply the document getting larger. A watermark on every page, page numbers, a signature image, an embedded font that was genuinely missing: these add bytes because they add content, and there is nothing to recover.
- Check the result rather than the number. A file that came back 70% smaller and 70% blurrier is not a success. Open it at 100%, read a paragraph, and if it is going to be printed read the small type; why is my PDF blurry covers what you are looking at if it has gone wrong.
Tools this article covers
Sources
Primary documentation for the claims above.