How it works

Why the page numbers in a PDF do not match the pages

Why don’t the page numbers in my PDF match the numbers printed on the pages?

Short answer

Because there are up to three numbers for the same sheet and nothing connects them. Its position in the file is what every tool counts and what a page range means. The label a reader shows in its page box is optional, and many applications never write one. The number printed on the page is ink — part of the picture, like any other mark — and nothing reads it. A document with a cover page and roman-numbered front matter has all three disagreeing by design.

Three numbers for the same sheet

The confusion is not a bug in anything. It is that the word “page” is doing three jobs at once.

351ii23

Three numbers for the same sheet

  1. Where the sheet sits in the file: the fifth of however many. This is what every tool counts, what a page range means, and it always exists.
  2. The label the document may carry for that sheet — roman numerals for front matter, a prefix for an appendix. Optional, and many applications never write one.
  3. The number printed on the page. It is ink: part of the picture, the same as any other mark, and nothing reads it.

Position in the file. The pages of a PDF are held in a tree, in order, and the fifth one is the fifth one. This number is not stored beside the page — it is simply the order — and it always exists, from 1 to however many pages there are. Everything mechanical uses it: page ranges, extraction, deletion, the number your reader scrolls past.

The label. A document may also carry a label for each sheet: i, ii, iii for front matter, then 1 onward for the body, then A-1 for an appendix. This is optional, it is a separate structure from the pages themselves, and a reader that supports it shows the label in its page box instead of the position.

The ink. The number the author printed in the footer is a mark on the page, drawn by the same machinery that drew the words above it. It is not a value the file records anywhere; it is a picture of a number. Nothing consults it, nothing can be wrong about it, and nothing will change it for you.

So a report with a cover, a contents page and two pages of front matter has a body that starts on sheet 5, may be labelled 1, and has 1 printed on it — three different answers, all correct, to “what page is this?”. The trouble only starts when two people, or a person and a program, mean different ones.

The label, and why your reader may not show it

Page labels are defined in ISO 32000-1 and they are stored more cheaply than you would guess. A document does not carry one label per page. It carries a small set of ranges, each saying: from this sheet onwards, number pages in this style, with this prefix, starting at this value.

1i2ii3iii4152637A-18A-2123

How page labels are actually stored

  1. Where each sheet sits in the file. This is the order of the pages, not a value written beside them, and it runs 1 to N whatever else the document says.
  2. The labels. A reader that supports them shows these in its page box instead of the position, which is how a document can open at a page called ii.
  3. They are stored as a few ranges, not one label per page: a style, a prefix and a starting number, applied from one sheet until the next range begins.

The styles available are decimal digits, upper- and lower-case roman numerals, and upper- and lower-case letters, each optionally with a prefix — which is where A-1 and Appendix 2 come from.

Storing ranges rather than values is efficient and has one consequence that catches everybody: the ranges are attached to positions, not to pages. Move a page and its label does not travel with it; the label belongs to wherever it now sits. Delete the first two sheets of a document whose labels restart at the third, and you have a document that begins at iii, or one with two pages both labelled 1.

Then there is reader support, which is patchier than the feature deserves. Some readers show labels in the page box and let you type them; others show the position regardless of what the file says; and the PDF Association’s survey of the field found that many viewers accept only arabic digits in the navigation box, so a document labelled ii and A-1 has pages you cannot navigate to by name even in software that displays the names. The same survey makes the more useful point for most people: the common word processors number pages perfectly well on the page and mostly do not export the label information at all.

Which means the ordinary case is a document with no labels whatsoever — its reader counts sheets, its footers say something else, and nothing is broken.

What a page range actually means

Every tool on this site counts sheets from the front of the file. Not labels, and certainly not printed numbers.

1-3147484950515223

A range box counts sheets, not printed numbers

  1. A page range, as every tool here reads it: positions in the file, counted from the front.
  2. So it takes the first three sheets, whatever happens to be printed on them.
  3. These are the printed numbers — ink, from whatever document these pages were cut out of. Asking for 47 to 49 asks for pages this file does not have.

So if you were given pages 47 to 52 of a contract as their own PDF, that document has six pages and they are pages 1 to 6. Typing 47-52 asks for something it does not have. This is the single most common way an extraction goes wrong, and the file is not to blame.

The two range boxes here behave differently from each other, which is worth knowing before you trust either:

  • Extract pages uses the shared parser. It accepts 1, 4-6, 12, takes -, an en dash or the word to between two numbers, and turns a reversed range the right way round. It then sorts and de-duplicates, so 3,1 gives you pages 1 and 3 in that order rather than in the order you typed. If you name a page that does not exist it refuses and says which page and how many there are, rather than guessing.
  • Split PDF’s page-list box is more forgiving and therefore quieter. It clamps a range that runs off the end — on a twelve-page document 1-500 gives you 1 to 12 — and silently ignores a single page number that does not exist. It keeps the order you typed. Its range mode is a set of from/to number fields instead, which is the version to use when it matters.

The habit that saves the afternoon: before typing a range, run the file through page count and file info and read the page count off it. It also reports the page dimensions and whether there is a text layer, which are the other two things people assume rather than check. How to split a PDF covers choosing between the modes.

What you can change, and what you cannot

Add page numbers to PDF prints numbers onto the pages. Precisely what it does, because the detail is the whole value of it:

  • It draws 10-point Helvetica, 28 units in from the edge, in one of five positions, in a dark grey. It uses one of the fourteen standard fonts, so no font data is added to the file — a few hundred bytes of drawing instructions per page and nothing else.
  • It has two separate start settings, and this is what people come for: which sheet to begin on, and what number to begin counting from. So a cover page and a contents page can go unnumbered while sheet 3 prints 1. Sheets before the start get nothing at all.
  • The n / total format counts the sheets being numbered, offset by the start number. Start on sheet 2 of a ten-sheet document and sheet 2 prints 1 / 9, not 1 / 10. That is arguably the right answer and it is certainly a surprising one, so it is worth seeing in the preview before you download.
  • Numbers are placed against each page’s own size, so a bundle mixing A4 and Letter — which merged documents routinely do — still gets its numbers in the same visual place on every sheet.
  • On a page carrying a rotation, the number is drawn into the page’s stored, unrotated space and the reader then turns the whole page, so on a sideways scan the number lands along what you see as an edge, reading sideways. The preview shows that rather than flattering it. Why a rotated PDF goes back to how it was explains the attribute behind it.

Two things it does not do, and nothing else here does either.

It does not write page labels. There is no tool on this site that sets them, so after adding numbers your reader’s page box still shows positions. If a filing requires the navigation box to match the printed numbers, that is a page-label job and it belongs in an application that can write them — ideally the one the document came from.

It does not remove numbers that are already there. Printed numbers are part of the page content; adding a second set gives you two. The only ways to take away printed ink are to cover it or to cut it off: redact paints over it and rasterises that page, and crop only hides what is outside the new boundary rather than deleting it. Both are worse than going back to the source and exporting again, if you can.

Where this bites

  • “See page 12” is ambiguous. In a document whose front matter is numbered separately, three sheets answer to it. Say which you mean — “the twelfth sheet” or “the page printed 12” — and the exchange ends there.
  • Searching for a printed number rarely works. In a scan it is pixels, so there is nothing to find. Even in a text document, a page number in a footer is frequently drawn separately from the body text and may not extract in a useful order. Why copied text comes out garbled covers why extraction order is a guess.
  • Reordering or deleting pages does not renumber anything. Printed numbers stay printed; labels stay with their ranges. And the tools here rebuild the document from copied pages, which starts from a fresh document catalogue — so a label scheme that was in the original is not carried into the output, and what you get back is plain positions. Reordering and deleting PDF pages covers what else moves and what does not.
  • Bookmarks and internal links point at pages, not numbers. They survive reordering because they reference the page object, which is usually what you want — and it also means a contents page whose printed numbers you have just invalidated will still take you to the right place while displaying the wrong number.
  • A merged bundle almost never adds up. Every source document brought its own printed footers, and the new document has one continuous position count. If the bundle needs to be citable, number it after merging and say in the covering note that the printed numbers in the body are the originals’.

Tools this article covers

Sources

Primary documentation for the claims above.

Last reviewed: September 30, 2026

Published by iBuildPDF.

More from the knowledge base

Browse the knowledge base