What makes a PDF accessible
What makes a PDF accessible to a screen reader?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Structure, not just text. A screen reader needs to know that this line is a heading, that block is a table with these columns, this picture means that, and the reading order runs this way. That information lives in a layer of tags that a PDF may or may not carry. Selectable text is necessary and nowhere near sufficient: a document can be perfectly searchable and still be unusable.
What tags are
Visually, a PDF page is a set of drawing instructions: put these glyphs at these coordinates. Nothing in that says which glyphs form a heading, or which belong to the same paragraph, or that these twelve runs of text are cells in a table. A sighted reader infers all of it from size, weight and position. Software cannot.
Reading order is stored, not guessed
- What a sighted reader sees: a heading, then two columns.
- The tag tree — heading, then paragraph, then paragraph. This is what says which column comes first.
- Without tags, the order has to be inferred from position, and a two-column page is where that goes wrong.
A tagged PDF adds a parallel structure tree — headings, paragraphs, lists, tables, figures — much like the structure of an HTML document. With it, assistive technology can announce "heading level 2", let someone jump between sections, read a table by row and column, and read a figure’s alternative text. Without it, the same software falls back to guessing from layout, and guesses badly on anything but the simplest single-column page.
Three further things the tag layer carries, all of which matter and none of which are visible:
- Reading order, which is not the order things were drawn. A two-column layout drawn column-first reads correctly; drawn line-first across both columns it reads as nonsense, and looks identical.
- Alternative text for images, without which a figure is announced as "graphic" and nothing more.
- The document language, which tells a screen reader which pronunciation rules to use. A French document read with English rules is close to incomprehensible.
Why OCR is not accessibility
This is the most common and most costly misunderstanding, so it is worth being blunt about.
Running OCR on a scanned document adds a layer of recognised text behind the image. The document becomes searchable, the text becomes selectable, and a screen reader will now read something rather than nothing. That is a real and large improvement, and it is not accessibility.
What OCR produces is an undifferentiated stream of words. It does not know that a line was a heading; it produces a line of text. It does not know a table is a table; it produces the cell contents in whatever order it scanned them, which for a table is frequently the wrong order and is announced without any indication of which column a number belongs to. It cannot write alternative text for a photograph, because it does not know what the photograph is of.
So: OCR is the necessary first step for a scan, and after it the document still has none of the structure that makes it usable. Treating "we ran OCR on it" as "we made it accessible" is how organisations end up believing a document archive is compliant when it is not.
Checking a document honestly
Some of this can be checked quickly, and it is worth doing before assuming a file is fine.
- Is there any text at all? Try to select a sentence. If nothing selects, the page is an image and needs OCR before anything else. Extract text from a PDF answers the same question — an empty result means there is no text layer.
- Is the reading order right? Select all the text on a complex page and paste it into a plain text editor. What you get is close to the order a screen reader will use. If a two-column page comes out interleaved, the reading order is wrong regardless of how it looks.
- Is it tagged? Desktop readers show this in the document properties, usually as "Tagged PDF: Yes/No". No is definitive; yes only means tags exist, not that they are correct.
- Is the language set? The metadata editor shows the document-level properties recorded in the file.
Automated checkers are useful and limited in a specific way: they verify that structures are present and well-formed, not that they are right. A checker cannot tell you that a heading was tagged as a paragraph, or that a photograph’s alternative text says "image1.jpg". Those need a person.
What browser tools can and cannot do about it
Being straightforward about the limits: iBuildPDF does not make PDFs accessible, and no tool that only edits an existing file really can. Accessibility is largely decided where the document is authored — a heading style applied in a word processor becomes a heading tag on export; text typed into a text box becomes an untagged fragment. Exporting correctly from the source is worth more than anything done afterwards.
What the tools here can do is the groundwork and the diagnosis:
- OCR gives a scan a text layer, which is the prerequisite for everything else.
- Extract text shows you what a machine actually sees, including the reading order.
- PDF to Word moves a document into an editor where headings, alternative text and table structure can be applied properly, and re-exported.
- The metadata editor exposes the document-level properties, including the title that assistive technology announces first.
Two operations actively make accessibility worse, and both are sometimes done for good reasons: converting pages to images destroys the text layer entirely, and flattening removes the interactive structure of form fields. Neither is wrong, but neither should be done to a document that someone needs to read with assistive technology.
If a document has a legal or contractual accessibility requirement, the standards to work to are PDF/UA (ISO 14289) and the PDF techniques published alongside WCAG, both linked below.
Tools this article covers
Sources
Primary documentation for the claims above.