Redaction fail with Epstein files

In December 2025, one of the biggest legal document releases in history exposed a mistake that professionals make every day. The fix is not a policy. It is a tool that does the job properly.

On 19 December 2025, the US Department of Justice released thousands of pages of documents connected to Jeffrey Epstein under the Epstein Files Transparency Act. Within days, social media users discovered they could highlight the black boxes covering sensitive names in certain documents, copy the text, and paste it into Word. The hidden content appeared in full. Files were quietly pulled from the DOJ website. Victims’ lawyers described an “unfolding emergency.” The failure was not a hack. It was a fundamental misunderstanding of how PDF redaction works.

The black box problem

When most people think about redacting a document, they picture drawing a black shape over the sensitive text. The result looks correct on screen. But a PDF contains two separate layers: the visual layer you see, and the text layer underneath that drives search, copy-paste, and every extraction tool on the market. A black shape affects only the visual layer. The text underneath is completely untouched.

This is not new. In 2019, lawyers for former Trump campaign chairman Paul Manafort filed a court document with black boxes drawn over sensitive passages. A Guardian reporter copied and pasted them into a new document and read that Manafort had shared presidential campaign polling data with an alleged Russian intelligence operative. In 2024, a journalist doing the same thing to “redacted” TikTok lawsuit documents found the company’s own internal research showing users can become addicted to the platform within 35 minutes.

Why AI makes the difference

There are two separate problems. The first is finding personal information in a document. The second is actually removing it.

On the first: a standard legal file might run 40 pages. Buried within it is a client’s ID number in a header, a cell number in a signature block, a medical aid number in an attached letter, an account number in a table. A human reviewer working under pressure will find most of them. Most is not good enough. AI finds all of them, in seconds. It does not get tired, does not scan past signature blocks, does not miss page 38. Every instance is flagged before a single redaction is applied and you stay in control of what gets removed.

On the second: when redaction is applied, it must permanently remove the underlying content from the file. Not cover it. Remove it. The text layer is gone. Copy-paste returns nothing. Search finds nothing. An AI tool extracting document text finds nothing because there is nothing left to find.

In 2026 this matters more than ever. Before AI became widely available, exploiting a failed redaction required knowing what to look for. Today, anyone can extract all text from a PDF with one instruction to ChatGPT or any other AI assistant. The copy-paste method is now the slow option.

It is happening here too

You might assume this is an American problem. South African court filings are not publicly searchable the way US federal documents are.

But the underlying failure, personal information believed to be handled, and not, is documented locally. Blouberg Municipality published a former employee’s personal details on its website in a declaration of interest and left them there even after warnings. The Information Regulator fined the municipality R500,000; a High Court confirmed the penalty in April 2026. Lancet Laboratories suffered multiple data breaches and failed to notify affected patients fined R100,000. The CIPC, whose systems every law firm uses for company searches — reported unauthorised access to client and employee data in February 2024. Your firm’s staff almost certainly had accounts there.

These are not redaction failures in the copy-paste sense. But the common thread is identical: someone believed a document or system was adequately protected. It was not. The Information Regulator reported 2,374 security compromise notifications in 2024/25, up 40 percent year on year and enforcement is accelerating.

SureDox

SureDox is a South African AI-powered redaction platform. Upload a document, a contract, a medical report, a court bundle, a discovery file and the AI scans it in seconds, detecting personal information across every relevant category: names, ID numbers, contact details, financial data, health information, addresses, company registration numbers. Items specific to the South African context , SA ID number formats, juristic person information that international tools miss are included.

Every detected item is presented to you for review. You decide what stays and what goes. When redaction is applied, the underlying content is permanently removed from the file. The document you download cannot be un-redacted by any method, because the data is no longer there. An audit log is attached to every output.

SureDox will not solve every compliance obligation your firm has. But it will ensure that the documents you send out do not carry personal information you intended to remove and that the removal is real, not painted. Try it at suredox.co.za, 10 free credits to test on your own documents, no credit card required.

The test worth doing today

Open a document your firm has previously sent out. Highlight the area you believe is redacted. Copy it. Paste it into a new document. If the text appears, the redaction failed and every document sent using the same method carried the same risk.

The Epstein files were handled by the Department of Justice of the United States. The Manafort documents were prepared by senior Washington DC attorneys. Blouberg Municipality had a legal team. Every one of them got this wrong. The question is whether your practice has a process that actually works or whether you are relying on the assumption that nobody will look.

References

Snopes: Can Epstein files be unredacted with copy and paste?
Confirmed copy-paste vulnerability in documents from the December 2025 DOJ release.

CBC News: How internet sleuths are un-redacting some of the Epstein files
Two un-redaction methods explained: copy-paste flaw and image exposure trick.

PBS NewsHour: At least 16 files disappear from DOJ Epstein site
Files removed from DOJ website shortly after the initial December 2025 release.

NPR: DOJ admits redaction errors in Epstein docs
DOJ confirms working to correct mistakes that left victim identities exposed.

Columbia Journalism Review: Manafort redaction failure
Guardian reporter copy-pastes Manafort filing and reveals Russia coordination details, January 2019.

NPR: TikTok lawsuit redaction failure
Kentucky journalist copy-pastes TikTok court filing, reveals internal addiction research.

ITWeb: InfoReg slaps DoJ with historic R5m fine
First POPIA administrative fine – 1,200 files compromised in 2021 DoJ ransomware attack.

ITWeb: InfoReg exposes POPIA violators as data breaches mount
2,374 breaches in 2024/25; 40% year-on-year increase; Blouberg and Lancet details.

Cliffe Dekker Hofmeyr: Recent privacy updates in South Africa
Blouberg R500,000 fine and Lancet R100,000 fine confirmed.

Ambledown: Regulatory Updates April 2026
High Court confirms Blouberg fine at R250,000 on appeal.

Adams & Adams: CIPC data breach
CIPC breach February 2024 – personal information of clients and employees exposed.

LEAVE A REPLY

Please enter your comment!
Please enter your name here

16 − 9 =