PDF basics

What PDF/A is, and why an archive rejects an ordinary PDF

What is PDF/A, and why was my PDF rejected?

Short answer

PDF/A is ordinary PDF with options taken away. Anything whose meaning could drift over time — a font the file only names, a colour defined by whatever screen you happened to use, encryption, scripts, a link to a file stored somewhere else — is forbidden, so the document has to be self-contained. An archive rejects a normal PDF not because anything is wrong with it today, but because it cannot promise the file will still render the same way in thirty years.

A PDF that is not allowed to depend on anything

PDF/A is defined by ISO 19005, which is a separate standard from the one that defines PDF itself. It adds no features. It removes options.

1A a B b234

PDF/A: a file that may not depend on anything

  1. The file. It is an ordinary PDF — any reader opens it, and nothing here is a new format.
  2. Everything it needs is inside it: the font outlines, a profile saying what the colours mean, and metadata describing the document.
  3. Nothing may be depended on from outside. That is the whole rule; the rest of the standard is that rule applied to each part of the file.
  4. So these are forbidden: encryption, scripts, references to content stored elsewhere — anything whose meaning could change after the file was written.

The problem it exists to solve is that an ordinary PDF is a file plus a set of assumptions about the machine that opens it: that the named font is installed, that this shade of red means something measurable, that the linked image is still on the server, that the reader can run the script. Each assumption holds until it does not, and the file gives no warning when one stops being true — it simply renders differently, and usually nobody notices which day that started.

PDF/A is the same format with those assumptions banned. Everything the document needs to draw itself must be inside the file.

Two consequences worth stating plainly. First, a PDF/A file is a PDF. There is no PDF/A reader and no PDF/A extension; every ordinary tool opens one, and most of the time you cannot tell by looking. Second, a file does not become PDF/A by being renamed, or by claiming to be one in its metadata. It either satisfies the rules or it does not, and the archive that rejected yours has a program that checks.

What it actually forbids

The rules are long, but almost all of them are the same rule applied to different parts of the file.

  • Every font must be embedded — including the fourteen that readers were once guaranteed to have. A file that only names a typeface is betting on a machine it will never meet. This is the single most common reason a document fails validation, and why PDFs look different on other computers covers what goes wrong when it is not done.
  • Colour must be device-independent. The file has to carry an output intent, or use colour spaces that are defined rather than assumed, so that a colour means a measurable colour instead of “whatever this monitor does”.
  • No encryption of any kind. An archive cannot preserve a file it cannot open, and a password is exactly the sort of thing that goes missing over decades. See what a PDF password actually protects.
  • No scripts, no launch actions, no embedded executables. Behaviour that depends on running code is behaviour nobody can guarantee later.
  • No external references. Everything drawn on the page must be in the file. A page that pulls an image or a font from a URL is a page that will eventually be blank.
  • XMP metadata is required, and it must agree with the document information fields rather than contradict them. What PDF metadata contains explains the two records and why they disagree so often.

A few rules depend on which part of the standard you are conforming to. Transparency and JPEG 2000 images are forbidden in PDF/A-1 and permitted from PDF/A-2 onwards; embedded attachments are forbidden in PDF/A-1, restricted to other PDF/A files in PDF/A-2, and unrestricted in PDF/A-3, which is what makes PDF/A-3 the usual carrier for an invoice with a machine-readable XML copy inside it.

One honest gap. A hyperlink to a website is allowed — it is an annotation, not content the page depends on — and it will still rot exactly as fast as any other link. PDF/A preserves the rendering of the document. It does not preserve the web.

The parts, and the letter after them

A conformance claim has two halves: which part of the standard, and which level within it. PDF/A-2b means part 2, level b.

The parts follow the underlying format. PDF/A-1 is built on PDF 1.4 and is the strictest. PDF/A-2 is built on ISO 32000-1 and relaxes the transparency and JPEG 2000 restrictions. PDF/A-3 is PDF/A-2 plus the ability to embed arbitrary files. PDF/A-4 is built on PDF 2.0 and reorganises the scheme — it drops the a/u/b letters, and adds a conformance level for interactive 3D engineering models, PDF/A-4e.

Within parts 1 to 3, the level says how much of the document’s meaning is preserved, and the levels are cumulative.

1U+004123

The conformance levels are cumulative

  1. Level b — basic. The file will look the same in future as it does now. That is all it promises.
  2. Level u — b, plus every character mapped to a Unicode value, so the text can still be searched and copied.
  3. Level a — u, plus a structure tree: headings, lists, reading order. The only level that cannot be reached by an export setting alone.
LevelWhat it guaranteesWhat it costs to produce
b — basicThe file will look the same in future as it does now.An export setting. Achievable for almost any document.
u — UnicodeLevel b, plus every character on the page maps to a Unicode value, so the text stays searchable and copyable.Usually an export setting too, if the producing application writes proper character maps.
a — accessibleLevel u, plus a structure tree: headings, lists, tables, reading order, language.Authoring work. This is not a checkbox.

PDF/A-1 has only levels a and b; the u level was introduced with part 2.

Most institutions that ask for PDF/A want PDF/A-1b or PDF/A-2b, and if you have not been told which, that is the pair worth asking about. Level a is a different order of work: a genuine structure tree has to come from a document that was written with headings and lists rather than bold text and tabs, and no converter reliably invents one. What makes a PDF accessible goes into why.

Where a PDF/A file has to come from

At the source. Word processors, layout applications, print pipelines and most report generators can export PDF/A directly, and that is the only route that reliably works, because the application still has the fonts, the colour information and the structure in front of it.

Converting an existing PDF to PDF/A is a repair job, and it is worth knowing what the repair involves before you trust the result. A converter has to find replacements for fonts that were never embedded — which means substituting something and accepting that the text may shift — assign a colour profile that was never chosen, remove encryption it may not be able to remove, and strip out interactivity. It frequently succeeds. It sometimes changes the document, and the changes are the kind nobody checks for.

iBuildPDF does not produce PDF/A. There is no tool here that converts a file to it, and a browser tool that claimed to without being able to locate a missing font would be claiming something it cannot do. What this site can do is tell you what you are holding before you go and find something that does:

  • Page count and file info — the structural facts, including whether there is a text layer at all.
  • The metadata editor — what the producing application recorded about itself, which is usually the fastest clue to how the file was made and therefore what it is likely to fail on.
  • PDF to text — a rough test of the u level. If the extracted text comes back as correct, readable characters, the character maps are there; if it comes back as nonsense, they are not, and the file cannot reach level u in its present state. Why copied text comes out garbled explains that failure.
  • Flatten PDF — resolves form fields and annotations into the page, which is one of the things a conversion would have to do anyway.

All four run in your browser. None of them makes a file conforming.

How to tell whether a file really is one

A PDF/A file declares itself in its XMP metadata, recording the part and the conformance level. Readers often show a bar across the top of the window saying so.

The declaration is a claim, not proof. Any file can carry it. A file can also be edited after it was made, so a document that was genuinely conforming on Tuesday may not be on Friday while still announcing that it is. The only way to know is to run a validator, and the archive that rejected your file ran one.

veraPDF is the open-source validator, developed with the PDF Association, and it is what a good many institutions use themselves. If your file has been refused, the useful thing to ask for is not another try but the validator report, which names the rule that failed. In practice most rejections are one of three things:

  1. A font is not embedded — often one font, often in a header or a page number nobody looked at.
  2. There is no output intent, so the colour is undefined.
  3. The file is encrypted, sometimes with a permissions password the sender did not know was there.

Last, the limit that matters most and gets said least: a file can pass validation and still be a poor archival document. A scan with no text layer is perfectly valid PDF/A-1b and will be unsearchable forever. Validation checks the container. Whether what is inside it is worth keeping, and whether anyone will be able to find it in twenty years, is a separate question that no standard answers for you.

Tools this article covers

Sources

Primary documentation for the claims above.

Last reviewed: September 28, 2026

Published by iBuildPDF.

More from the knowledge base

Browse the knowledge base