
What counts as personal data in an AI prompt
You open a chat with an AI tool, paste in a work question, and hit send. It feels harmless. It almost never is. The text you just sent often carries more about real people than you realised, and once it leaves your screen you no longer control where it sits or who can read it later.
Most people picture personal data as a passport number or a credit card. Those count, but the category is far wider than that. The good news is that once you can spot it, you can deal with it. You do not have to stop using AI. You just have to remove the personal parts first.
What personal data actually means
The clearest definition comes from the GDPR, the European data protection law. Article 4 defines personal data as any information relating to an identified or identifiable natural person.
Read that slowly. Two words do the heavy lifting. "Identified" means you already know who the person is. "Identifiable" means you could work out who they are, directly or by piecing other details together. So personal data is not only the obvious identifiers. It is anything that points back to a specific human being.
The law also marks out special categories that get extra protection. These include health information, biometric data, and other sensitive details about a person. If your prompt touches any of those, you are handling the most protected kind of data there is.
The two ways data points to a person
It helps to split identifiers into two types.
Direct identifiers name the person on their own. A full name, an email address, a phone number, a national ID number, a passport number, or a home address each point to one individual with no extra work needed.
Indirect identifiers are the quiet ones. Each looks harmless alone, but combine a few and they single someone out. A job title plus a company plus a city can be enough. A birth date plus a postcode can be enough. None of these is a name, yet together they describe exactly one person.
This is why "I removed the name, so it is fine" is a trap. The name is often the easiest part to strip and the least of your worries. The combination of small facts is what quietly gives someone away.
What hides in an everyday prompt
Here is a normal work request, the kind people send dozens of times a week:
"Draft a firm but polite email to Maria Gonzalez at maria.g@northbridge.com about her overdue invoice INV-20471 for the consulting work. She mentioned she has been off sick recovering from surgery, so keep the tone gentle. Her direct line is +34 612 445 980 if she wants to call."
Look at how much personal data is packed into three sentences:
- A full name, which is a direct identifier
- An email address, another direct identifier
- A phone number, a third direct identifier
- An invoice number and account context that tie to a real client relationship
- A health detail, recovering from surgery, which falls into the special category that gets extra protection
That last one matters most. A health reference about a named person is exactly the kind of sensitive data the rules treat with the most care, and it slipped in almost by accident, just to set the tone of the email.
Now imagine that same prompt sent to a tool you do not control, stored on a server you have never seen, possibly read by people you will never meet. You wanted help writing an email. You handed over a small dossier on Maria.
Why this is easy to miss
Personal data hides because it is doing a useful job in your sentence. You include the name so the email feels real. You mention the surgery so the tone lands right. You add the phone number so the reply is easy. Every detail is there for a sensible reason, which is exactly why your eye slides past it.
The other reason is habit. When you talk to a colleague you trust, you share context freely, and that is fine because you both work for the same employer under the same rules. An AI tool can feel like that trusted colleague. It is not one. It is a system that may keep, move, or process what you send in ways you cannot see.
The fix: remove the personal parts first
The takeaway is simple. The goal is not to use AI less. The goal is to send it the problem without the people.
In the example above, the AI does not need to know the client is called Maria Gonzalez. It does not need her email, her phone number, or her medical history. It needs to know that a named client has an overdue invoice and a sensitive personal reason for a gentle tone. Swap the real details for safe placeholders and the AI writes you the same quality email, because the shape of the request is unchanged. The personal parts were never what made the answer good.
The important word is reversibly. If you blank out the details by hand, you get back a draft full of gaps you have to fill in again, and people quickly give up and just paste the raw version. What you want is a way to mask the personal parts on the way out and restore them on the way back, so the final result is ready to use with the real names back in place.
That is the whole idea. Spot the personal data, direct and indirect, including the sensitive categories. Replace it before it reaches the AI. Get a useful answer. Put the real details back yourself, on your side, where they belong.
Velum does exactly this. It masks the personal parts of a prompt before they ever reach an AI model, then restores them in the response, so your team keeps the speed of AI without handing over the people behind the words. To see it on your own prompts, request a demo.