Why a table from a PDF comes out wrong in Excel
Why does my table come out as a mess when I convert a PDF to Excel?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Because the table was never in the file. A PDF records each piece of text at a position on the page, and the ruling lines as separate drawing instructions that are not attached to anything. No row, column or cell is stored. A converter has to work out where the columns were from how the values line up, so a merged heading, a value that wrapped onto two lines, or a table with no lines around it can each move the boundaries and shift the data into the wrong cells. Nothing is corrupt when that happens; the information the converter needed was not recorded in the first place.
There is no table in the file
This is the part that is worth believing before anything else, because every other surprise follows from it.
The same page, as a reader sees it and as the file holds it
- What you see: rows, columns and cells, with lines making the structure obvious.
- What the file records: each value as its own piece of text at its own position on the page. Nothing joins the three values of a row to each other.
- The ruling lines are drawing instructions, like any other mark on the page. They are not attached to the text and a table drawn without them is stored identically.
When a word processor or a spreadsheet exports a PDF, it draws the table. It writes INV-1001 at one position, 2026-01-04 at another, 1,240.00 at a third, and then draws some lines. The lines are drawing instructions like any other — the same machinery that drew a logo or an underline. Nothing in the file says that those three values belong to one row, or that the line between two of them is a column boundary rather than a decorative rule.
It is easy to test. Take a page with two tables on it, one with ruling lines and one without, and ask any text extractor what is on the page. Both come back as the same thing: a list of values, one after another, in roughly the order they were drawn. The ruled table and the unruled one are indistinguishable in the output, because the rules were never part of the data.
A PDF can carry real structure — tagged PDF records which marks are a table, a row and a cell, mostly for screen readers. It is optional, most exports do not produce it, and most of the ones that do produce it imperfectly. What makes a PDF accessible covers what tagging is and why so little of it exists in the wild.
So when a converter hands you a spreadsheet, it did not read a table out of the document. It looked at where the text landed and made a case.
What the converter is actually guessing
Two things, and the second is much harder than the first.
Where column boundaries come from, and what moves them
- The boundaries are inferred, not read: a stripe of whitespace running the height of the table is taken to be the edge of a column.
- A value long enough to reach across that stripe closes it, and the boundary that depended on it moves or disappears.
- Text that wrapped inside its cell sits on a new baseline, so it is read as a new row with every other cell empty.
Rows are usually easy: text sitting on roughly the same baseline is one row. This mostly works, and it is why a simple table usually converts well.
Columns have to be inferred from the vertical gaps — the gutters that run down the page where no text ever appears. Find a stripe of whitespace that is clear from the top of the table to the bottom, and that is a column boundary. It is a good heuristic. These are the things that break it:
- A value long enough to close the gutter. One long description or one large number with separators can reach across the stripe that defined the boundary, and every row below may be read as having one column fewer.
- Text that wrapped inside a cell. The second line sits on its own baseline, so it is read as its own row — and the cells beside it are empty, which is how a 40-row table arrives with 60 rows in it.
- Merged or spanning headings. A title centred over three columns lines up with none of them, and can drag the header row out of alignment with everything beneath it.
- Right-aligned numbers beside left-aligned text. The gutter is not straight, so its edges are fuzzy and the boundary lands in a different place depending on which rows are sampled.
- A table that continues onto the next page. Each page is reconstructed on its own evidence, so a repeated header becomes a row of data in the middle, and column boundaries found on page two need not match page one.
None of this is a defect in a particular converter. It is a reconstruction problem, and reconstructions are sometimes wrong.
What PDF to Excel on this site does
PDF to Excel is a table extractor. That is not a detail of how it is built, it is what the tool is, and knowing it saves a lot of confusion:
- It looks for tabular structure and writes what it finds into a workbook. It is not trying to reproduce the page.
- Prose is not a table, so prose does not come out. A page of paragraphs produces nothing, and that is the correct answer rather than a failure.
- It runs on this site’s processing server, not in your browser — this is one of the jobs a browser genuinely cannot do. The file is uploaded, converted and deleted after the job; the notice on the tool page states what happens to it.
Two related tools are often the better choice:
- PDF to Word tries to rebuild the page — layout, headings, and tables as Word tables. If what you want is the document rather than the numbers, start there. Converting PDF to Word without losing formatting covers what survives.
- Extract text runs in your browser and gives you the raw text with no reconstruction at all. For a small table, pasting that into a spreadsheet and splitting it yourself is often faster than arguing with a converter, and you can see exactly what you are working with.
When it says there was nothing to extract
This message means the converter ran, read the pages and found no tabular structure. It is a real answer, and there are three usual reasons for it.
The page is a picture. A scan, a photograph, or a screenshot pasted into a document has no text in it at all — just pixels that look like text. There is nothing for an extractor to position, so there is nothing to find. Run OCR first to add a text layer, then convert. Is my PDF scanned or text? shows how to tell in a few seconds, and OCR or PDF to Word? explains which you need.
What you are looking at is not a table. A list with aligned numbers, a form, or a two-column layout can look like a table to a reader and present no table-like structure to an extractor. Extract text is the right tool for those.
The table is laid out loosely. Wide, uneven spacing with no consistent gutters gives the heuristic nothing to lock onto. Converting that page alone sometimes works where converting the whole document did not, because the evidence is no longer being averaged across pages that disagree.
Getting a better result
- Go back to the source if it exists. If the PDF came from a spreadsheet, that spreadsheet still has the real rows and columns. Every conversion is a reconstruction; the original is not.
- Convert the pages with the table, not the whole document. Use extract pages first. Fewer pages means less conflicting evidence, and you find out quickly whether the problem is one awkward page.
- Check the result against the PDF before you use it. Compare the row count and add up one column. Quiet errors — a shifted column, a row that merged into the one above — are the ones that cause damage, because nothing looks broken.
- If the layout is hopeless, extract the text instead. Extract text plus Excel’s own text-to-columns gives you control over the split, and for a table under a few dozen rows it is usually quicker.
- Scanned pages need OCR, and OCR of a table is doubly a guess — first what each character is, then where the columns were. Check figures individually; a 6 read as an 8 does not look wrong anywhere.
- Check the page count first. Page count and file info also reports whether the document has a text layer at all, which answers the scanned question before you convert anything.
Tools this article covers
Sources
Primary documentation for the claims above.