Converting and extracting

Saving a web page as a PDF

Why does a web page saved as PDF come out badly?

Short answer

Because a web page has no pages. It is one continuous column that adapts to the window, and making a PDF means cutting it into fixed sheets at positions nobody chose. Headings land at the bottom of a page, tables split down the middle, and anything positioned relative to the window ends up somewhere unintended.

Two formats with opposite assumptions

An HTML page is written not to have a fixed size. It reflows for a phone or a wide monitor, it can load more content as you scroll, and its layout is calculated at the moment you look at it.

A PDF is the opposite in every respect: fixed page size, fixed positions, decided once. See what a PDF actually is.

So conversion has to invent the information that the web page never had — how wide the page is, and where each sheet ends. The usual casualties:

  • Sticky headers and floating buttons repeat on every page, or sit on top of the text.
  • Cookie banners and newsletter pop-ups are captured along with everything else.
  • Anything that loads as you scroll may be missing, because it was never loaded.
  • Wide tables are cut off at the right edge rather than wrapped.
  • Dark themes print as dark, using a great deal of ink.

Where the cuts land

The converter fills a sheet, cuts, and starts the next one. It has no idea that the line it just cut through was a heading.

123

Cutting a continuous page into sheets

  1. A web page: one column with no pages in it and no natural place to stop.
  2. Where the cuts land — decided by how much fits, not by what the content means.
  3. The result. A print stylesheet is what stops a cut landing in the middle of a heading or a table.

Well-built sites include a print stylesheet — instructions for exactly this situation, telling the browser to hide navigation, avoid breaking inside a table row, and keep a heading with the text beneath it. When a page converts beautifully, that is why. When it converts badly, the site simply never considered printing, and there is nothing in the page to guide the cutting.

HTML to PDF does the conversion and honours a print stylesheet where one exists.

Getting a clean result

  • Look for a print or reader view first. Many articles have one, and it is a page built for exactly this — no navigation, no sidebar, no banner.
  • Scroll to the bottom before converting, so anything that loads lazily has loaded.
  • Dismiss the banners before you capture, or they become permanent.
  • Use landscape for wide tables. It is the one setting that reliably rescues them.
  • Trim afterwards rather than fighting the layout. Crop PDF removes a repeated header margin across every page in one pass, and Remove pages drops the comments section.
  • Combine related pages into one document with Merge PDF — useful for a documentation set saved page by page.

If the captured page is enormous because of full-width photographs, Compress PDF will usually take most of it back.

What cannot be captured

Anything behind a login that a converter cannot reach. A tool fetching a URL sees the page as an anonymous visitor, so a personal dashboard or a paywalled article comes back as a sign-in screen. Printing from your own browser, where you are signed in, is the way round that.

Anything interactive. Maps, charts that respond to hovering, video, forms, and anything driven by a script become a still picture of their current state — or of their empty state, if they had not finished drawing.

Pages that block automated fetching. Some sites refuse requests that do not come from a browser, and will return an error page instead of their content.

And one point worth stating plainly: capturing a page produces a copy of someone else’s work. Saving an article to read offline is ordinarily fine; redistributing it is a different matter, and that is a question about permission rather than about file formats.

Tools this article covers

Sources

Primary documentation for the claims above.

Last reviewed: September 24, 2026

Published by iBuildPDF.

More from the knowledge base

Browse the knowledge base