Why text copied out of a PDF comes out garbled
Why does text I copy from a PDF paste as nonsense?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Because what a PDF shows and what it can hand you are two separate maps. The page draws numbered codes through a font, which is what makes it look right; copying needs a second table saying which character each code stands for, and that table is optional. When it is missing or wrong the page is flawless and the clipboard is gibberish, and no amount of retrying changes it.
A PDF draws glyphs, not characters
The content stream of a page does not contain your text. It contains numbers.
Two maps, and only one of them is required
- The page stores codes, not letters. On their own these numbers mean nothing.
- The font turns each code into an outline. This is the map a PDF must have, and the one that makes the page look right.
- Copying needs the other map — code to Unicode character — and a file may be written without it. Then there is nothing to hand over.
- So the page can be flawless and the clipboard nonsense at the same time. Nothing is broken; a table was never written.
Each number is a code. The font resource attached to the page maps that code to an outline — a shape to draw. That is the only map a PDF is required to have, because it is the only one needed to put marks on a page, and it is why the document looks perfect.
Copying needs the map going the other way: this code means the character U+0041, capital A. ISO 32000-1 provides for exactly that, as a ToUnicode CMap attached to the font. It is optional. A producer that only cared about the document printing correctly had no reason to write one, and nothing in the file complains about its absence.
Subsetting makes the failure worse. When a producer embeds only the glyphs a document actually uses, it is free to number them however it likes — commonly 1, 2, 3 in the order the characters were first encountered. Without a ToUnicode map those numbers mean nothing whatsoever; with one, they mean whatever it says. Why PDFs look different on other computers covers subsetting and the other surprises it causes.
The consequence worth absorbing: this is a property of the file, not of your computer. Changing reader, operating system, font settings or paste target will not help, because they are all consulting the same missing table. It is also why searching inside the document finds nothing — the reader is comparing what you typed against the same broken map.
Which of these is yours
The symptom identifies the cause fairly reliably, and the cause decides whether anything can be done.
| What you get | What it is | Can it be fixed |
|---|---|---|
| Nothing selects; the cursor does not catch on any word | There is no text. The page is an image. | Yes, by recognising it — OCR |
| Every character wrong, consistently — a shifted alphabet, or rows of boxes | Missing or incorrect ToUnicode map | Not in place. OCR, or retype |
Mostly right, but fi, fl and ffi vanish or become one odd character | Ligature glyphs with no character mapping | Yes — find and replace afterwards |
Right characters, no spaces: thequickbrown | The file contains no space characters at all | Often, with a different extractor |
Right characters, spaces inside words: t h e | Wide letter-spacing in the layout being read as gaps | Often, with a different extractor |
| Right words, wrong order — two columns interleaved, a footnote mid-sentence | Extraction is following the order the marks were written | Sometimes. Depends on the file |
| Accents detached, or letters splitting into two characters | Combining marks stored separately from their base letters | Yes — Unicode normalisation |
The first row is the one to rule out before anything else, because it is common and it is not this article’s problem. If nothing highlights when you drag across the page, you have a scan, and there is no text to garble.
The same page, with and without text
- A digital PDF. On screen, indistinguishable from the one on the right.
- Underneath it: character codes, each with a font and a position. Searchable, selectable, convertible.
- A scanned PDF. The same page, and the same appearance.
- Underneath it: a grid of light and dark samples. No letters, no words, nothing to search.
Is my PDF scanned or text covers telling the two apart in a couple of seconds.
Spaces and order are a different problem
These two produce readable-looking wreckage rather than nonsense, and they have a different cause from the missing map.
There may be no spaces in the file
Nothing requires a PDF to contain space characters. A line of text is typically drawn as a series of strings with numeric offsets between them, and the offset is what makes the gap between two words. Whether that gap is a space is a judgement an extractor makes from the width: too generous and words run together, too eager and words split in half. This is a heuristic, which is why two tools disagree on the same file and neither is wrong.
The order is the order it was drawn in
Why two columns come out interleaved
- Two columns, drawn correctly. You see columns because of where the marks are, not because the file says so.
- The order the marks were actually written in. A layout engine often alternates between frames, and nothing records that these were columns.
- An extractor follows that order, so the columns come out interleaved line by line. A tagged PDF states the real reading order; an untagged one has nothing to consult.
The content stream is in whatever order the producing application emitted the marks. A layout engine working through text frames may alternate between two columns; a template may place the footer before the body; a table may be drawn cell by cell down the columns rather than across the rows. Nothing in an ordinary PDF records that these marks were a column, a footnote or a table — the columns exist in your reading of the page, not in the file.
Extractors compensate by sorting marks by position, which handles simple two-column pages well and complicated layouts badly. A tagged PDF is the exception: it carries a structure tree stating the real reading order, and an extractor that honours it gets the right answer without guessing. That is one of the genuine practical benefits of tagging, quite apart from the accessibility requirement it usually exists for — see what makes a PDF accessible.
What actually helps, in order
- Try a different extractor before anything else. PDF to text runs in your browser on PDF.js and honours
ToUnicodemaps where they exist. If it returns correct text, the file was never the problem and your reader’s copy behaviour was — which is a two-minute discovery worth making before a two-hour one. - If the map is broken, no text-layer tool will help. Every extractor is reading the same table. Converting to Word, to a spreadsheet or to plain text all consult it, so they all produce the same nonsense in different wrappers. This is the point at which to stop trying tools.
- Ignore the text layer and read the page as a picture. That is OCR: render the page, recognise the shapes, write a fresh text layer. It is the only reliable route out of a broken map, and it has two costs — recognition makes its own errors, and the tool runs on a server, so the document is uploaded. For a confidential file that is a decision rather than a click.
- If it is spacing or order, extract and clean up. A find-and-replace pass over the extracted text usually fixes ligatures and run-together words faster than any tool. For a document you need to keep working on, PDF to Word gives you something editable to repair — also server-side, and what it can and cannot recover is worth reading first.
- If you are the one producing the file, fix it at the source. Export rather than print to PDF, embed the fonts, and tag the document. A file exported from a current word processor almost always copies correctly, and the ones that do not are usually the ones that went through a printer driver on the way.
And the unglamorous one: for a page or two, retyping is often faster than any of the above. It is worth saying out loud, because people spend an hour on a five-minute job out of a feeling that the computer ought to be able to do it.
What no tool can do
- It cannot reconstruct a map that was never written. The information is not hidden or encrypted, it is absent. Anything that claims to fix a broken
ToUnicodemap is inferring characters from glyph shapes, which is optical character recognition with a different label on it. - OCR gives you a new best guess, not the original. Its errors cluster on exactly what you can least afford to get wrong — digits, reference numbers, surnames, anything you cannot check by reading for sense. Proofread anything you are going to act on.
- Correct extraction is still not the document. Tables come back as lines, footnotes come back in the middle, captions come back adrift from their figures. Getting every character right tells you nothing about structure, and no amount of character-level accuracy ever will.
- Right-to-left and complex scripts have a second failure above this one. Some producers write Arabic or Hebrew in visual order, using presentation forms rather than the underlying letters. A correct character map then yields text that is technically right and reads backwards, and putting it back into logical order is its own job.
The useful habit is to test before you commit: take one page, extract it, and read what comes out. Two minutes at the start decides whether this document is a copy-and-paste job or a retyping job, and that is the only decision that actually matters here.
Tools this article covers
Sources
Primary documentation for the claims above.