How to redact a PDF so the text is actually gone
A black rectangle in a PDF is a drawn object, not a deletion — the text underneath still copies out. What actually removes it, and how to check before you send.
The single most common way people leak data in a PDF is also the most confident-looking: draw a black rectangle over the sensitive line, save, send. It looks finished. In most tools it is not.
This is the same mistake as blurring text in a screenshot, but it fails harder and faster. A blur at least requires arithmetic to undo. A black box over PDF text requires selecting the text and pressing Ctrl+C.
Why the black box does not work
A PDF is not a picture of a page. It is a program that draws a page.
The file holds a content stream — a sequence of drawing operators. Text is stored as operators that say set this font, move to this position, show these glyphs. When you add a rectangle in most editors, you append one more operator: fill this area black.
Painting happens in order, so the rectangle lands on top and you stop seeing the text. Nothing removed it. The glyphs and their coordinates are still in the stream, sitting below the rectangle in paint order and completely intact in the file.
Which means all of the following still work on a “redacted” PDF:
- Select and copy. Drag across the black box in any reader and paste. The text comes out because the text layer is still there and selection reads the layer, not the paint.
- Text extraction.
pdftotext, or any library that walks the content stream, returns the full text and never renders anything at all. The rectangle is invisible to it. - Search. Ctrl+F finds words under the box and highlights them, sometimes drawing the highlight over the rectangle.
- Moving the object. In an editor, the rectangle is a separate object. Click, drag, and the text is underneath.
None of this is exotic. The first two are what a journalist, an opposing lawyer or a curious recipient tries within a minute of receiving a document that visibly has something hidden in it.
What actually removes it
Real redaction rewrites the content stream: it deletes the text operators covering the selected region and then draws the box. Two steps, and the order matters — the deletion is the part that counts, the box is only so a reader can see that something was removed.
Acrobat Pro does this correctly, and the wording in its interface is the thing to watch. Marking a redaction only marks it. The file is not changed until you use Apply redactions, which rewrites the page and cannot be undone afterwards. A file that was marked but never applied looks identical on screen to one that was, and is not redacted at all.
Follow it with Sanitize document, which is a separate operation and removes the parts that are not page content — metadata, attachments, form field values, scripts, hidden layers.
Other tools vary, and the interface will not tell you which kind you have. A feature called “redact” in a free web tool or a general PDF editor is often just a filled shape with a confident label. Do not infer behaviour from the button name. Test it with the check below, once, on a file you do not mind ruining, and then you know for that tool forever.
The check that settles it
Thirty seconds, no special software, and it answers the only question that matters.
- Save and close the file, then reopen the saved copy — not the editor’s live document.
- Drag-select across the redacted area, from clear text on one side to clear text on the other.
- Paste into a plain text editor.
If the hidden words appear in the paste, the redaction is decorative and the file is not safe to send. If you get a gap, the text operators are gone.
Two extra passes for anything going out publicly:
- Ctrl+F for a word you removed. Search reads the text layer directly and will find it even when selection behaves oddly.
- Extract the text on the command line.
pdftotext file.pdf -prints everything the file will give up to anything automated. This is the check that matches what an adversary would actually run, and it takes one command.
The failure that survives correct redaction
Even a properly applied redaction can leak through the file’s history.
PDFs support incremental updates: saving an edit can append the changes to the end of the file and leave the previous version in place above it. The reader shows you the newest revision. The older one — including the page as it was before you redacted it — can still be sitting in the same file, recoverable by anything that walks the earlier cross-reference tables.
This is why “Save As” to a new file, rather than “Save”, is the safer habit, and why Acrobat’s Save As Optimized or a sanitize pass is worth doing on anything sensitive: both rewrite the file rather than appending to it.
The other quiet leaks are the ones that are not page content at all:
- Document properties — author, company, the original filename, the software used.
- Attachments and embedded files, which readers show in a side panel most people never open.
- Form fields. A flattened-looking form can still carry its field values as data.
- Comments and annotations, including ones set to not display.
- Optional content groups, PDF’s layers. A layer can be switched off and still be in the file.
- The filename.
settlement-jane-smith-confidential.pdfundoes the work inside it.
The route that always works, and its cost
If you do not have a tool you trust, there is a fallback that cannot fail for the reason described above: render the page to an image, then redact the image.
Rasterising the page throws away the text layer entirely. There are no glyph operators left to copy out, because there are no operators at all — just pixels. Then redact the image the way you would redact any image, with an opaque fill rather than a blur, and rebuild a PDF from the results if you need one.
Be precise about the order. Rasterise, then redact. If you draw the box first and then rasterise, you are fine too — rendering only writes what is visible — but if you keep the un-rasterised file anywhere, you have kept an unredacted document.
The cost is real and worth stating plainly:
- The document stops being searchable and selectable, for you as well as for everyone else.
- Screen readers can no longer read it. For anything public-facing or legally required to be accessible, this alone rules the approach out.
- File size usually grows, sometimes a lot.
- Text quality drops at low render resolutions.
For a one-page invoice going to one person, that is an easy trade. For a 200-page filing, it is the wrong answer and you should get a tool that redacts properly.
Where this overlaps with what we build
Be clear about the boundary first, because it decides whether any of this is useful to you. SnapShield AI does not open PDFs. There is no import path for them. If the job in front of you is redacting an existing PDF — a filing, a contract, a report someone sent you — the tools earlier in this article are the answer and nothing here replaces them.
What it does do is the other end: it captures a page and exports PDF, and the way it builds that file is the reason it is worth mentioning in an article about text surviving redaction.
The export takes the flattened image and wraps it in a single-page document. In the code
that is one DrawImage call onto an empty page — the whole page is that image. There are
no text-showing operators written into the content stream, because no text is ever placed
there. Which means the failure this entire article is about cannot occur in a file it
produces: there is no text layer under the redaction to select, copy, or extract with
pdftotext, because there is no text layer at all.
That is worth being honest about in both directions. It is safe by construction rather than by care, which is a genuinely stronger guarantee than “we remembered to apply the redaction”. It is also the rasterised route described above, so it inherits every cost listed there: the resulting PDF is not searchable, not selectable, and not readable by a screen reader. For a support ticket or a one-page record, that is the right trade. For a document that has to be accessible or searchable, it is not, and you should redact the original properly instead.
The redaction itself is the part that has to hold up, and it is the same requirement as for any image: blur is reversible, so an image route is only safer than a black box if the redaction genuinely replaces pixels rather than filtering them. Ours does, and the exported file is what the tests assert against.
Two practical notes rather than claims: PDF export is a Pro feature, and capture, annotation and export to PNG are on the free tier, so the redaction behaviour can be checked without paying for anything.
In short
- A black rectangle is paint, not deletion. The text is still in the file.
- Marking a redaction is not applying it. Check which one you did.
- Verify on the saved file by selecting, pasting, and searching.
pdftotextif it matters. - Save As, not Save, so an incremental update does not carry the original along.
- Rasterising works when nothing else is available, at a genuine cost in accessibility and searchability.
Related reading
- Can blurred text be recovered? — the same failure in image form, with the maths.
- How to redact a screenshot on Windows — the tool-by-tool version once your page is an image.
- Redacting screenshots for GDPR — what has to come out when the data belongs to someone else.
- Watermarking screenshots
Try it on your own screenshot
SnapShield AI redacts on your machine — no upload, no account, no expiry. Free tier available, 84 MB, Windows 10 and 11.
Download Free