Imagine this: You’re using a powerful Large Language Model (LLM) like ChatGPT to analyse a lengthy contract or summarise a complex case. You paste the text, eager for the time-saving insights. But what you might not realise is that you’re not just sharing the text with the AI; you’re often sharing the underlying data structure, metadata, and potentially years of embedded information. This is the unseen witness, and it poses a significant, often invisible, risk when handling sensitive client data.
In the age of AI, upholding this duty requires more than just good intentions; it demands rigorous data hygiene.
The Sneaky Perils of “Hidden” Data
When we think of redacting a document, we often picture blacking out text with a marker. In the digital world, people mistakenly replicate this by using the drawing tool in Word to place a black box over sensitive text. This is a critical security flaw.
Here’s why:
- The Text is Still There: Most PDF viewers and word processors treat a “redaction box” as an annotation, not a structural change to the underlying text. An AI (or even a person with basic technical knowledge) can simply select and copy the text behind the black box.
- Metadata is the Smoking Gun: Your document contains hidden information about its creation and editing history. This includes the author’s name, the date and time of creation, comments, tracked changes (which can reveal previous drafts), and even GPS coordinates if the document was created on a mobile device. AI models can, and do, process this metadata.
- Contextual Clues: Even if you think you’ve scrubbed the names, an LLM might infer identity. For example, a uniquely descriptive sentence combined with a general location might be enough for the AI (or a subsequent user querying the AI) to re-identify the subject.
- Training the AI LLMs: The fundamental risk is why the AI is using your data. Many public LLMs use your queries to further train and improve their models. If you paste a confidential client’s private financial data, that data could theoretically influence the AI’s future responses to other users. While many enterprise-level AI solutions offer “no-training” clauses, public models are less strict.
In short, simply “hiding” or “covering” the data isn’t enough. It must be permanently and digitally removed from the file’s code.
Join the Conversation: AI for Lawyers Fireside Chat
We understand that navigating these technical nuances can be daunting while you’re already managing complex legal work. That’s why we’re inviting you to join our upcoming AI for Lawyers Fireside Chat.
Our core topic will be “How to Remove Sensitive Data” (properly!).
The Details:
- Topic: AI for Lawyers Fireside Chat – How to Remove Sensitive Data
- When: This coming Wednesday at 4:00 – 4:30 PM
- Where: Zoom – at your desk with a fresh cup of coffee
- How: By invitation to subscribed members of AI for Lawyers Fireside Chat
Ready to Make Your Practice AI-Safe?
The potential of AI for law is immense, but it must be met with equally robust security. Join us this Wednesday to ensure you’re not just using AI, but using it safely.
Register as a member of our Fireside Chats here: https://tech4law.aweb.page/exclusive-ai-lawyer-chats
We look forward to seeing you there.








