Redacting a folder of screenshots: how to do it at scale without missing one

Manual redaction does not scale past about twenty images, and it fails silently. How a batch pass works, what it misses, and how to verify the output.

· 8 min read · Nexoradia Labs

One screenshot you can check properly. Twenty you can check if you are disciplined. Two hundred — a support ticket export, a QA run, a year of bug reports being moved to a public tracker — you cannot, and the reason is not laziness.

Attention on a repetitive visual task degrades in a specific way: you stop reading the image and start recognising its shape. By image sixty you are confirming that a screenshot looks like the ones before it, not checking whether this one has an email address in the top-right corner. The mistakes cluster at the end of the batch and they are invisible to the person making them.

So a folder of screenshots needs a different method from a single one.

What a batch pass actually does

The mechanics are unglamorous, and that is the point: an input folder, an output folder, and one detection-and-redaction pass per image, written out as a copy.

For each supported file — PNG, JPG, JPEG, BMP, WebP, TIFF, GIF — the pass decodes the image, runs the detectors over it, applies the chosen redaction to every region that came back, and writes a redacted copy to the output folder. Originals are read, never modified. Failures are recorded per file rather than aborting the run, so one corrupt image in the middle of two hundred does not cost you the other 199.

Two things about that design matter more than they look:

Separate output folder, always. In-place redaction is a category of accident, not a feature. If the detection was wrong you have destroyed the only copy that could tell you so, and if the run crashes halfway you now have a folder in two different states with no way to tell which files are which. Read from one directory, write to another, and diff the counts at the end.

A per-file failure list. A batch that reports “done” without telling you it skipped eleven files is worse than one that fails loudly. Check the count in equals the count out before you ship anything.

In SnapShield AI this runs off the UI thread with a progress report per file, and can be pointed at faces, text, or both. The important part for you is what it detects and, much more importantly, what it does not.

What detection catches

Two independent passes run over each image and their results are merged.

The first is model-based: face detection, plus Windows OCR to find text regions. The second is structural pattern matching over the recognised text — and structural is the operative word. Credit card numbers are validated with the Luhn checksum, IBANs with the mod-97 check, so a 16-digit order reference is not blacked out as a card number and a real card number is not missed because it was formatted with spaces. Alongside those: email addresses, phone numbers, JWTs, API keys in bare, bearer, labelled and vendor-prefixed forms, national identifiers, IP addresses, usernames, passwords, bank details, addresses and personal names.

That covers the cases that are both the most common and the most tedious. It is a good first pass. It is not a review.

What it does not catch — read this part twice

Every limitation below is a way a batch run can report success on an image that still leaks.

  • OCR is the ceiling. If the OCR pass cannot read the text, the pattern matching never sees it. Low contrast, unusual fonts, heavy JPEG artefacts, a photo of a screen, a screenshot of a screenshot — all of these read as “no text found”, which looks exactly like a clean image in the results.
  • A missing Windows language pack looks like a clean batch. If the relevant language pack is not installed, text detection quietly finds nothing across every file. Two hundred images come back reporting nothing sensitive. Before a real run, test on one image you know contains an email address. If that comes back empty, fix your Windows language settings before trusting anything else.
  • Faces at an angle, partly occluded, small, or badly lit are missed by face detectors in general. This one is no exception.
  • Context is not detected at all. A licence plate, a name badge, a whiteboard, an internal hostname in the URL bar, a door number, a wallpaper photo of someone’s children, a taskbar showing which applications a person had open — none of these are faces or pattern matches. They are yours to catch.
  • Sensitive by domain, not by shape. A diagnosis, a salary, a disciplinary note, an internal project codename, a customer’s company name — all ordinary text. No pattern matcher knows they matter in your context.

The honest framing is that a batch pass removes the tedium of the obvious cases so your attention is free for the non-obvious ones. It does not replace looking at the images.

A verification strategy that fits the volume

Verifying every one of two hundred outputs defeats the purpose. Verifying none is how leaks happen. Sample by risk instead:

  1. Every image the run reported zero detections for. This is the highest-value bucket by a wide margin. A zero is either a genuinely clean screenshot or a detection that never ran — and those two look identical in a results table. There are usually few enough to check individually.
  2. Every image that failed or was skipped. These have no redacted output at all. It is remarkably easy to ship the originals for these by accident when you copy the leftovers across.
  3. A random ten percent of the rest, opened at full size, not as thumbnails.
  4. Anything from a source you know is unusual — a different app, a different display scale, a screenshot someone else took, an image that came in over chat and has been recompressed.

For each image you check: open it in a different program from the one that made it, zoom to 400% on each redacted region, and push brightness and contrast to the extremes. Solid regions stay flat. Anything that ghosts under those adjustments was never removed.

Choose the redaction style before the run, not after

A batch is where the wrong choice does the most damage, because you make it once and it applies to everything.

Use solid blocks for anything that must not be recoverable. Blur and pixelation across a folder of screenshots is worse than across one, because a batch gives an attacker something a single image does not: many samples of the same font, the same UI, the same blur radius, and often the same values in different places. That is a much easier reconstruction problem than a single blurred region — here is what recovery actually involves.

Blur is fine when the goal is tidiness rather than protection, and that distinction is worth being honest with yourself about before you start a run of two hundred.

After the run

Deal with the leftovers on disk. Auto-save folders, crash-recovery files, capture histories, thumbnail caches, and the input folder itself all still contain the unredacted originals. Your output folder is clean; your machine is not. Decide deliberately whether the originals are archived somewhere access-controlled or deleted, and do that before the output folder gets shared.

Then check the counts one last time: files in, files out, files failed. If those three do not reconcile, the batch is not finished.

Try it on your own screenshot

SnapShield AI redacts on your machine — no upload, no account, no expiry. Free tier available, 84 MB, Windows 10 and 11.

Download Free

Keep reading