Key takeaways
- A rule that is not memorable at six on a deadline is not a rule, it is a document.
- Four categories is the most a person will retain. Any more and they will guess instead.
- Your AI provider is a subprocessor in substance, and most client agreements have something to say about that.
- Personal accounts remove your ability to respond to an incident. That is the strongest argument for consolidation.
Most firms write their AI data policy the same way they write every policy: comprehensively, with definitions, categories and a decision tree. It is thorough, it is defensible on paper, and it fails at the only moment that matters, which is a Thursday evening when someone needs to summarise a long document to meet a deadline and has to decide in four seconds whether that is allowed.
Policies that are not recallable under time pressure do not govern behaviour. They govern the aftermath.
Write for the moment of decision
The practical test for any rule in this area is whether a mid-level professional, three months after the training session, under deadline pressure, can recall it and apply it without opening anything. That standard rules out almost every policy document written on this subject.
It admits four categories. Not because four is a natural number of categories for data, but because four is roughly what a person retains.
| Category | Rule | Why |
|---|---|---|
| Public or already published | No restriction | Nothing to protect |
| Ordinary internal and client working material | Firm account only | Covered by your agreement, excluded from training, retrievable |
| Restricted: personal data of individuals, health, payment, credentials | Ask before it goes anywhere | Specific legal regimes attach and the answer varies |
| Never: material under a named-recipient confidentiality agreement, or anything a client has restricted | Does not go in at all | You have already promised otherwise in writing |
The second row carries most of the daily volume and it is the one people get wrong in both directions. Some firms forbid it, which drives usage onto personal accounts and makes the situation worse. Others say nothing, which leaves people guessing. Stating plainly that ordinary client working material belongs in the firm account and nowhere else is the single most useful sentence in the whole policy.
The one-line version that people actually remember: if it is client work, it goes in the firm account and nowhere else; if it is personal data, health, payment or credentials, ask first; if a client or an agreement has told us to restrict it, it does not go in. Three clauses. Put it on the wall, not in a binder.
The subprocessor problem most firms have already
This is the part that surprises people, and it is worth checking before it surprises you in a client audit.
If an AI provider processes client material on your behalf, it is functionally a subprocessor, whatever your internal terminology. A large number of client agreements, master services agreements and data processing addenda contain provisions about subprocessors: a requirement to maintain a list, to give notice before adding one, to obtain approval, or to flow down specific terms.
Firms that adopted AI tools without revisiting those agreements are frequently in breach of commitments they signed years before the technology existed. Nobody has noticed yet because nobody has asked, which is not the same as being compliant.
The remediation is straightforward and worth an afternoon:
- Pull the data protection and subprocessor clauses from your top clients by revenue.
- Identify which require notice, a list, or approval.
- Add your AI provider to your subprocessor list, which many firms discover they do not maintain at all.
- Give notice where required, framed as routine housekeeping rather than as a disclosure. It is a normal update and treating it as one sets the tone.
Doing this proactively is a considerably better experience than doing it in response to a question from a client's procurement team.
What actually happens to the data
Staff and clients both ask this, and a firm should be able to answer it in plain terms rather than by referring to a policy.
On a business or enterprise tier, the position with the major providers is generally that your inputs are not used to train their models, that content is retained for a defined period to provide the service and to meet abuse monitoring obligations, that provider staff access is restricted and logged, and that you can configure retention within limits. On a consumer tier, several of these change, and the training exclusion may be a user-toggled setting rather than a contractual commitment.
That distinction is the entire argument for consolidation and it is worth stating in exactly those terms, because it converts an abstract governance point into something a professional can evaluate.
Keep the evidence, not just the assurance. Screenshot the training and retention settings the day you configure them, record the date and who did it, and re-check quarterly. When a client asks, a dated screenshot is a materially better answer than a paragraph describing what you believe to be true.
The questions people actually ask
Does removing names make it safe? Usually not. Removing an obvious identifier does not remove identifiability, particularly in a small market where a description of the transaction identifies the parties to anyone in the industry. Treat de-identification as a specialist exercise rather than an editing habit.
What about material that is confidential but not personal data? Commercially sensitive material is governed by your confidentiality obligations rather than by privacy law, which means the constraint comes from the agreement rather than from a statute. Read what you promised. Some agreements are more restrictive than any privacy regime.
Where is the data processed? Ask, and get the answer in writing, because some clients and some regimes care and the answer varies by provider and tier. For most US mid-market firms serving US clients this is not a live issue, and for any firm with European or government-adjacent clients it very much is.
Can we use client material to improve our own internal tools? Careful. Building an internal knowledge base out of client work is a different act from using a tool to process it, and most engagement terms did not contemplate it. This one needs a specific decision and usually a specific permission.
When something goes into the wrong place
It will happen. Having a defined response in advance is what separates a contained incident from a bad month.
The sequence: establish what was disclosed, when, and into which account. Delete what can be deleted, including the conversation. Check the contractual position on retention and training for that account tier. Assess whether it meets any notification threshold under your client agreements, your professional obligations or applicable law. Record what happened and what you did, because the record is what demonstrates you took it seriously.
The decisive variable is the account. On a firm account with a business agreement you have a contractual position, a retention setting, deletion capability and a vendor you can contact. On someone's personal account you have almost none of that, and the disclosure was made to a party you have no agreement with.
That difference, more than any policy language, is the argument for getting everyone onto one governed account.
Frequently asked questions
What client data can safely go into an AI tool?
On a business tier with training exclusion in the contract, ordinary client working material is generally defensible, subject to your engagement terms and any client-specific restrictions. Personal data, health, payment and credential material needs a specific check first, and anything under a named-recipient confidentiality agreement should not go in at all.
Is our AI provider a subprocessor?
In substance yes, if it processes client material on your behalf. Many client agreements require subprocessor notice, a maintained list or approval. Firms that adopted AI without revisiting those clauses are often in technical breach of agreements signed years earlier, and proactively giving notice is far better than being asked.
Is deleting names enough?
Usually not. Removing an obvious identifier does not remove identifiability, especially in a small market where the facts of a matter identify the parties to anyone in the industry. Treat de-identification as a specialist exercise rather than an editing habit.
What do we do if something goes in that should not have?
Establish what, when and into which account. Delete what can be deleted, check the retention and training position for that tier, assess notification obligations, and record it. The account matters most: a firm account gives you contractual recourse and deletion capability, a personal one gives you almost nothing.
Get the data questions settled once
We configure the account so the answers are true, evidence the settings, set retention and access, cap the spend, and write the short rules your people will actually follow along with the subprocessor position your clients ask about. Fixed price, live in 30 days or less.