Governance & compliance

What is PII in AI projects?

Feeding customer data into AI tools without controls is a common and serious exposure. Where data goes, how long it is kept, and who can see it must be settled before launch.

In plain terms

Personally identifiable information is any data that can identify a specific person, and data privacy is the practice of protecting it throughout an AI system.

Also called PII.

PII is personally identifiable information: anything that can point to a person. Names, emails, phone numbers, account numbers, and combinations of smaller details that still identify someone. Health and payment data sit in stricter buckets on top of that.

AI projects touch PII when staff paste a ticket into a chatbot, when a transcript is stored, or when a model is trained on customer mail. The rule is simple enough to brief a team. Do not put personal data into a tool whose contract does not allow it. Remove or mask it when the task does not need it. Know where the text goes and how long it stays.

Think of it this way: Feeding customer data into an AI tool without controls is like leaving a customer file open on a shared desk. The exposure is not intentional; it is structural, and that does not reduce the liability.

A company uses a consumer AI tool to summarize customer support transcripts. The tool sends those transcripts to a third-party model. When a GDPR review surfaces this, legal halts the tool immediately.

A support agent pastes a full customer email, address included, into a personal ChatGPT account to draft a reply. The draft is fine. The address is now outside the company. The allowed tool masks the address before the model sees it, and personal chat accounts are turned off for that team.

Map every data flow in your AI system before launch: what data enters, what model processes it, where it is stored, and who can access it. This is standard practice, not a compliance formality. There is no context where personal data can flow through an AI system without a legal basis, retention policy, and access control. These requirements do not scale with project size.

RaftLabs writes the control into the system: which data can enter, who approves the result, and how you explain it later. The rule and the product stay the same story. The related work on our side is AI governance.

This sits with the other governance & compliance terms on the glossary. The rules that keep AI legal, and keep customer data out of the wrong tool. Worth reading next: GDPR & Compliance, Model Governance, and Shadow AI.

Common questions

Only under a contract that says what the vendor may do with it, where it is stored, and whether they train on it. Consumer chat accounts are not that contract. If the task does not need the name or the account number, strip them first. When in doubt, ask privacy before the pilot, not after the paste.
A combination that still identifies someone. An employee ID, a rare job title plus a location, a device identifier, a free-text complaint that names a street and a date. If you would not post it on a noticeboard, do not paste it into an unapproved tool.

Work with us

Tell us what's broken.

Tell us what's not working in your business. We'll find the real problem and tell you exactly what it would take to fix it.

  • Scope and cost agreed before work starts. No surprises. No obligation.
  • Working prototype within 3 weeks of kickoff.
  • Pay by milestone. You see progress before each invoice.
  • 60-day post-launch warranty. Bug fixes, UI tweaks, and deployment support. No retainer.
  • All conversations are NDA-protected.