How to redact a PDF so the text cannot be recovered
How do I black out text in a PDF so it genuinely cannot be recovered?
Save your favourite tools
Create a free iBuildPDF account to keep your favourite tools and find them quickly anytime.
Free account. The PDF tools themselves never need one.
Short answer
Drawing a black rectangle over text does not remove the text — the words are still in the file underneath and come straight back out with copy-paste. Real redaction means destroying the underlying content, which is what /redact-pdf does by rebuilding the marked pages as images, and you should always verify the result by trying to extract the text again.
The short version
A PDF is a list of drawing instructions. Adding a black rectangle over a name appends one more instruction: draw a filled black rectangle at these coordinates. The instruction that draws the name is still in the file, earlier in the same list. The rectangle sits on top of it visually and does nothing to it structurally.
Select the area in a PDF reader and press copy. The name is on your clipboard; nothing was cracked, the text was simply read out of the file the way any reader reads any text.
Redaction that holds up means removing content, not hiding it. Everything else here follows from that distinction.
Why black boxes fail
Under ISO 32000-1, the standard that defines PDF, the marks on a page come from a content stream: a sequence of operators that position a font, set a colour, and show a string of characters. Text is stored as text. That is what makes a PDF searchable, and why a reader can hand you the words when you select them.
Covering is not removing
- A black rectangle drawn on top of the text. It looks finished.
- The text is still in the file, underneath. Select it, copy it, or open the file in any editor and there it is.
- Proper redaction rebuilds the page without those characters in it.
- Now there is nothing beneath the mark, because there is no longer anything to find.
A shape drawn afterwards is just another operator in the same stream. Painting order decides what you see; it does not decide what exists. So the words survive every route into the file that does not involve looking at it:
- Select and copy in any PDF reader.
pdftotext, or any other extraction utility, which ignores graphics entirely and returns the text layer.- Search — the document is still indexable on the hidden words, including by search engines if it is published.
- Opening the file in an editor and deleting the rectangle, which restores the original appearance exactly.
This failure mode has produced real disclosures. The pattern recurs in court filings and government releases: a document is published with names or figures covered by black rectangles, and within hours someone has copied the text straight out. The US National Security Agency publishes guidance on sanitising PDF files for exactly this reason, and its central instruction is the one above — covering content is not removing it.
The same applies to every variation. Highlighter yellow over black text, white text on a white background, a coloured box, or a shape in an annotation layer are all cosmetic changes to what is drawn; the characters remain in the content stream and a text extractor returns them. White-on-white is the worst of the set, because it does not even look redacted.
What actually works
Proper redaction has one requirement: the content that carried the sensitive information must no longer be in the file. There are two ways to reach that.
The first is surgical — parse the content stream, remove the text-showing operators covering the redacted region, and rewrite the stream. Precise, but fragile: text is positioned in fragments that rarely align with words. Done imperfectly it leaves partial strings behind, which is the worst outcome available, because the file then looks redacted and is not.
The second is to rebuild the page so the original text objects no longer exist at all. This is what iBuildPDF’s redaction tool does. You mark the regions to remove; for each marked page, the tool renders that page with PDF.js, draws your redaction areas as solid fill onto the rendered pixels, and writes a new page built from that image with pdf-lib. The rebuilt page contains one image and no text operators. There is nothing underneath the black areas because there is nothing underneath anything — the page is now pixels.
The trade-off is real and you should plan for it. On every page you redact, all text stops being selectable, not just the redacted part. Those pages are no longer searchable, copyable, or readable by a screen reader, and the file size usually goes up, because an image of a page of text is larger than the text was.
Pages you do not mark are left exactly as they were, so a 200-page report with two redacted pages keeps 198 searchable ones. Where the trade-off is unacceptable — an accessibility requirement, or a document that must stay searchable throughout — redact at the source instead: remove the text in the original and export a fresh PDF. Either way, redact last: re-exporting from the original afterwards silently reinstates everything you removed.
Checking that it worked
Never ship a redacted document you have not tested. The test takes two minutes.
- Copy and paste. Open the downloaded file, select across the redacted area and a generous margin around it, copy, and paste into a plain text editor. On a correctly rebuilt page you get nothing, because there is no text to select.
- Extract the whole document. Run the result through PDF to text and search the output for the redacted string. This is the check that matters most, because it covers the entire file rather than the area you happened to select — including a page where the same name appears and you forgot to mark it. Search fragments too: a surname alone, the last six digits of an account number.
- Check the metadata. Open the file in the metadata editor. Sensitive terms routinely survive in the Title or Keywords fields, and the Title is very often the original filename —
Smith-settlement-confidential.pdfdefeats the redaction on page four without anyone reading the page.
All three tools run in your browser, so verifying a confidential file does not mean uploading it. If a search does turn something up, treat the file as unredacted rather than patching it with another rectangle.
What redaction cannot do
Four limits, stated plainly.
- It only affects the copy you download. Redaction produces a new file. The original is unchanged on your disk, and any copy already sent, filed or backed up still contains everything. If the unredacted version has left your control, redacting it now changes nothing about that.
- Order matters when rasterising. Rebuilding a page as an image is only safe because the redacted area is filled in before the image is written — the black pixels replace the pixels that showed the text. A process that rasterises the page first and then places a black rectangle on top of the resulting image has recreated the original problem in a new form: the words are still there as pixels under the box, recoverable by anyone who removes it. Confirm with the extraction test above rather than assuming a tool got the order right.
- A page image can still be read by a machine. Rasterising defeats copy-paste and text extraction; it does not make the remaining visible text unreadable to optical character recognition, an ordinary capability of widely available software. Redaction protects the covered content, not the rest of the page. iBuildPDF does not provide OCR — there is no such tool on the site — so this is about what someone else can do with your output.
- Metadata is a separate system. The document information fields and the XMP block live outside the page content and are untouched by anything done to pages. Clear them with remove PDF metadata, and see what PDF metadata contains for what those fields hold.
The habit that prevents most redaction failures is simple: assume the file still contains what you removed, and make it prove otherwise by extracting the text and searching it.
Tools this article covers
Sources
Primary documentation for the claims above.