What personal information is hidden in a PDF
What personal information is hidden inside a PDF, and how do I remove it?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
A PDF carries a set of hidden fields — most commonly an author name taken from your operating system account, the name and version of the software that made the file, and creation and modification timestamps — none of which you typed and none of which appear on the page. They can be read by anyone who opens the file and removed with a metadata stripper such as /remove-pdf-metadata.
The short version
Open any PDF you have sent to someone and look at its properties. You will usually find a name in the Author field. You did not type it. It came from the account name on the computer that produced the file — which, on a work laptop, is often your full legal name or an internal username, and on a home machine is whatever you called yourself when you set the thing up.
That field is one of several. A PDF is a structured document format, and the specification (ISO 32000-1, the standardised form of Adobe’s PDF 1.7) sets aside dedicated places for information about the document that is not part of what is drawn on the page. Word processors, scanners, print drivers and export dialogs all fill those places in, generally without telling you.
None of this is an exploit; it is the format working as designed. It matters because the design assumes you know it is happening, and most people sending a CV or an anonymous complaint do not.
What is actually in there
There are two separate metadata systems in a modern PDF, and a file can carry both, often with different values.
The part of a PDF nobody reads
- What is on the page. This is the part you checked before sending it.
- The record attached to it: author, software, dates, sometimes a file path or a company name.
- None of it appears when the document is printed or read, which is exactly why it travels unnoticed.
The document information dictionary
This is the older mechanism, defined in ISO 32000-1. It is a small set of named fields attached to the document:
| Field | What it usually contains |
|---|---|
Title | Often the original filename, including a draft name or an internal project code |
Author | Almost always the operating system account name of whoever created the file |
Subject | Usually empty, sometimes a description typed years ago into a template |
Keywords | Usually empty; when populated, frequently reveals internal categorisation |
Creator | The application the content was authored in — Word, InDesign, LaTeX, a scanner’s driver |
Producer | The library that wrote the PDF itself, normally including its exact version number |
CreationDate | Timestamp of creation, including the time zone offset |
ModDate | Timestamp of the last modification |
Two deserve particular attention. Author is the field that identifies a person, and it is filled in automatically. Producer identifies software down to the version, which tells a reader what you run and, by implication, what your organisation deploys.
The timestamps are informative too. A time zone offset places you geographically, and a ModDate after a document was supposedly finalised is the kind of detail that gets noticed in disputes.
XMP metadata
The newer mechanism is XMP — the Extensible Metadata Platform, standardised as ISO 16684-1. It is an XML block embedded in the file, and it holds more than the eight fields above: rights statements, editing history, tool-specific records, and a document identifier that persists across saves and can link separate files back to a common ancestor.
Because XMP is separate from the information dictionary, the two can disagree. Editing one does not necessarily update the other, which is how a file ends up showing a scrubbed Author in its properties panel while the original name sits in the XMP block a few thousand bytes away.
Metadata inside embedded images
A PDF containing photographs embeds those images as objects within the file. If an image carried EXIF data when it was placed — camera model, capture timestamp, and on phone photos very often GPS coordinates — that data can travel into the PDF with it. This is a third, independent layer, untouched by anything that edits the document’s own metadata.
How to see what your file is carrying
The fastest way is the PDF metadata editor. Drop a file in and it parses the document and shows the information dictionary fields with their current values. The file is read in your browser and not uploaded, so you can inspect a confidential document without handing it to anyone.
Check the file you are about to send, not the one you started from: every export step rewrites Producer, and some rewrite more.
The page count tool is a useful companion for a quick structural look — it opens the document and reports how many pages it contains, which also confirms the file parses cleanly as a PDF at all. A file that will not report a page count is damaged, and repairing it comes first.
Outside the browser, pdfinfo (part of Poppler) and exiftool both dump metadata, and exiftool also reports EXIF found on embedded images. A PDF is partly plain text, so opening one in a text editor and searching for /Producer or <x:xmpmeta will often show the fields directly.
Removing it
Use remove PDF metadata. It reads the document with pdf-lib, clears the document information fields — Title, Author, Subject, Keywords, Creator, Producer, and the creation and modification dates — and writes out a new PDF. The page content itself is not rewritten, so the document you get back looks and reads exactly like the one you put in.
If you would rather replace values than empty them, the metadata editor lets you set each field explicitly. That has a genuine use: an organisation name in the Author field is less conspicuous than a blank one, which in a batch of otherwise normal documents is itself a signal.
Both tools run entirely in your browser. Nothing is uploaded, which matters more than usual here — the whole point is that the file contains something you do not want to share, so sending it to a stranger’s server to have that removed is a strange trade. See how browser PDF tools work for the mechanics.
One practical note: do this as the last step before sending. Anything that rewrites the file afterwards — re-exporting, printing to PDF, opening and saving in another application — will stamp a fresh Producer and new timestamps back in.
What stripping metadata does not do
This is the part that gets people into trouble, because clearing the properties panel feels conclusive and is not.
- It does not touch the visible page. A name in a letterhead, a footer reading Draft — J. Okonkwo, a header carrying a client’s name: these are page content, not metadata. They are still there, and they are the first thing a reader sees.
- It does not remove EXIF inside embedded images. The document’s metadata and an embedded photograph’s metadata are different structures. Clearing the first leaves the second intact, GPS coordinates included. If a PDF contains phone photos and location matters, strip the EXIF from the images before you build the PDF.
- It does not remove text hidden under a drawn shape. A black rectangle placed over a paragraph does not delete the paragraph — the text objects are still in the content stream and come straight back out with copy-paste. That is a separate and much more serious failure, covered in how to redact a PDF permanently, and it is not solved by any metadata tool.
- It does not reach copies already sent. Metadata removal produces a clean new file. Every copy distributed before that still carries everything, in every inbox and backup it reached. Cleaning is done before a file leaves, not after.
Treat metadata as four separate layers: the information dictionary, XMP, embedded image EXIF, and visible page content. Each needs its own check. Clearing one and assuming the rest followed is the mistake worth avoiding.
Tools this article covers
Sources
Primary documentation for the claims above.