PDF basics

What a PDF file actually is

What is a PDF, underneath?

Short answer

A PDF is a description of finished pages. It records exactly where every glyph, line and image sits on a page of a fixed size, along with the fonts needed to draw them. That is why it looks the same everywhere — and why editing one is awkward. There are no paragraphs inside a PDF, only marks in positions.

What is inside the file

A PDF is four parts in a row, and three of them are readable if you open the file in a plain text editor.

%PDF-1.712xref3trailer4

The four parts of a PDF file

  1. Header — one line naming the PDF version the file uses.
  2. Body — the numbered objects: pages, fonts, images, and the instructions that say what to draw.
  3. Cross-reference table — the byte position of every object, so a reader can jump straight to one.
  4. Trailer — where the table starts and which object is the root of the document. Readers start here, at the end.

The header is one line naming the version. The body is a numbered collection of objects — pages, fonts, images, and the content streams that say what to draw where. The cross-reference table records the byte offset of every one of those objects. The trailer says where that table begins and which object is the root of the document.

The table is the part worth understanding, because it explains most of a PDF’s behaviour. A PDF is not read from beginning to end like a letter; it is looked up. A viewer opens the file, jumps to the end, reads the trailer, finds the table, and from then on goes straight to whichever object it needs. That is how a reader shows you page 400 of a 900-page report without reading the first 399.

It is also why a file that lost its last few kilobytes in transit usually will not open at all. The objects are still there; the map to them is not. Why a PDF will not open covers that case, and Repair PDF can sometimes rebuild the table from what survives.

To see a file’s own account of itself, Page count and file info reports the structural facts and the metadata editor shows the descriptive ones.

Why it looks the same everywhere

Every position in a PDF is absolute. A line of text is not “the second paragraph”; it is a set of glyphs placed at stated coordinates on a page of a stated size. Nothing reflows when the window changes, because there is nothing to reflow. The layout decisions were made once, by whatever produced the file, and then frozen.

Fonts are the other half of the promise. A PDF can carry the fonts it uses inside itself, so a reader draws the text with the same shapes the author saw rather than substituting whatever the machine happens to have. When a font is not embedded, the substitution shows — different letter widths, different line breaks, occasionally a different alphabet. That is nearly always the cause of a document that looks wrong on someone else’s computer, and it has its own article.

This fixedness is the whole value of the format. It is also the whole cost of it, and the two cannot be separated: a document that is guaranteed to look identical everywhere is a document that cannot adapt to anything.

Versions, and the PDF/A question

The version in the header — 1.4, 1.7, 2.0 — says which features the file may use, not how new or good it is. Readers are backward compatible in practice, so a 1.4 file from 2003 opens fine today. You almost never need to care, and “save as an older PDF version” is rarely the fix for anything.

The variants are worth knowing about:

  • PDF/A is for archiving. It requires every font to be embedded and forbids anything whose meaning could drift — no external references, no scripts, no encryption. A PDF/A file is a bet that it will still render correctly in thirty years.
  • PDF/X is for commercial printing, and pins down colour and page geometry. See preparing a PDF for printing.
  • Tagged PDF adds a structure tree — headings, lists, reading order — which is what makes a document usable with a screen reader. What makes a PDF accessible goes into it.

These are constraints layered on the same format, not different formats. A PDF/A file is a PDF, and every ordinary tool can read it.

Why editing a PDF is awkward

Because there is no text to edit, in the sense a word processor means. Change “March” to “September” and the replacement is longer, so it collides with whatever came next — and nothing in the file knows that those glyphs formed a sentence, or that the sentence belonged to a paragraph that should re-wrap. A PDF editor has to reconstruct that structure by inference, and it is inference, which is why results vary with how the file was made.

What this means in practice:

  • Small corrections — a date, a number, a name — are what a PDF editor is genuinely good at.
  • Rewriting is better done at the source. Edit the original document and produce a fresh PDF; if the original is gone, PDF to Word gets you an editable approximation, and what it can and cannot recover is worth reading first.
  • Page-level work — reordering, deleting, rotating, merging — is easy and lossless, because it moves whole objects rather than touching their contents. Reordering and deleting pages explains why.

One warning that matters more than the rest: covering text with a black rectangle in an editor hides nothing. The text is still in the file, underneath. Redacting a PDF properly is a different operation, and Redact PDF is the tool for it.

Tools this article covers

Sources

Primary documentation for the claims above.

Last reviewed: September 24, 2026

Published by iBuildPDF.

More from the knowledge base

Browse the knowledge base