Key takeaways
- "Is our data safe" is four separate questions. Answering them separately makes the problem tractable.
- Training exclusion, retention, processing location and internal access are all different, and only one of them is what people usually mean.
- Your client contracts constrain you more often than your AI provider does.
- Get the answers in writing at the tier you are buying, and keep the document for due diligence.
Ask a board whether they are comfortable putting company information into an AI tool and you will get a reasonable hesitation without a specific objection. That vagueness is the actual problem. Vague concerns cannot be resolved, so they become a permanent handbrake, while the organisation carries on pasting the same material into consumer tools without any of the protections.
Breaking the worry into its four real components makes each one answerable in an afternoon.
Question 1: is our content used to train models
This is what most people mean and it has the clearest answer. At business and enterprise tiers, the major providers generally exclude customer content from model training by default. Consumer and free tiers frequently do not, and that difference is one of the main reasons the business tier exists.
Two practical notes. Confirm it for the specific tier you are purchasing, because the terms differ between tiers of the same product. And take the wording from the contract or the official documentation rather than a sales conversation, then save a copy, because you will need to quote it on client questionnaires for years.
Worth internalising: even where content is excluded from training, it is still being transmitted to and processed by a third party. Training exclusion answers one question well. It does not answer the other three.
Question 2: how long is it kept
Separate from training, and often more relevant to your actual obligations. Providers retain conversation history so users can return to it, and retention periods vary by provider and tier. Business tiers usually let you configure this.
The decision involves a genuine tension. Shorter retention reduces how much of your material sits in someone else's system, which is good for confidentiality and for breach exposure. Longer retention preserves work people rely on, and may be required by record-keeping rules in your sector.
There is no universal answer, but there is a wrong one, which is accepting the default without knowing what it is. Decide it, document the reason, and revisit annually.
Question 3: where is it processed
This is the question with actual legal weight in many jurisdictions and sectors. Where does processing physically occur, and which entity is the counterparty on your contract?
It matters most if you operate under data localisation requirements, hold government or health sector contracts, or have clients whose agreements specify where their information may be handled. If none of those apply, this is a question to answer and file rather than one to agonise over.
Ask specifically: which regions handle processing, whether region can be constrained at your tier, which legal entity you are contracting with, and what the arrangement is for onward transfers to subprocessors.
Question 4: who at the provider can see it
Providers generally have limited internal access for abuse investigation, security incidents and support, with controls around it. This is normal for business software and is the same arrangement you already accept from your email, storage and CRM vendors.
It is worth asking about because the answer usually reassures, and because it puts AI in the same category as the other systems you already trust rather than in a special category of its own. Consistency in how you assess vendors is itself a governance improvement.
The constraint people forget: your own contracts
In professional services, the binding limitation is usually not the AI provider's terms. It is what you have already promised your clients.
Many client agreements restrict which third parties may process their information, require notification before adding a subprocessor, or contain confidentiality wording drafted long before AI tools existed and broad enough to cover them. These clauses do not care that the third party is a chat interface rather than an offshore team.
The practical sequence:
- Pull your standard client agreement and read the confidentiality and subprocessor clauses with AI specifically in mind.
- Identify which clients have bespoke terms that differ from the standard, because those are where the surprises live.
- Decide whether your default position is to ask, to notify, or to rely on existing wording, and apply it consistently.
- Add AI-aware wording to new engagement terms so this stops being a retrospective problem.
An underrated move: tell clients proactively that you use AI tools under a governed configuration, and say which controls you have. Increasingly this reads as competence rather than risk, and it is far better than the same fact emerging during a dispute.
What to ask a provider, in order
| Question | What a good answer looks like |
|---|---|
| Is our content excluded from model training at our tier? | Yes, by default, with a clause reference you can quote. |
| What is the default retention period, and can we change it? | A specific period, and an admin setting you can see. |
| Where is processing performed, and can we constrain the region? | Named regions, and a clear statement about whether constraint is available at your tier. |
| Who at your organisation can access our content, and under what circumstances? | A narrow, documented set of circumstances with access controls described. |
| What happens to our data if we cancel? | A defined deletion timeframe, in writing. |
| Which certifications do you hold, and can we see the report? | Current certifications and a report available under a confidentiality agreement. |
| Do connectors respect our existing permissions? | Yes, with an explanation of how, which you should then test yourself. |
Collect these answers once, keep them in a single document with dates, and reuse it. Most organisations answer the same client due diligence questions repeatedly and rebuild the answer each time.
The controls that do most of the work
Once the questions are answered, the protective measures are unremarkable and effective:
- Buy the business tier, so the default terms are the ones you actually want.
- Set retention deliberately rather than inheriting it.
- Restrict external sharing of conversations to inside the organisation.
- Connect single sign-on so access ends when employment does.
- Give people a short, concrete list of what may and may not go into the tool.
- Eliminate personal accounts for work use, which is where nearly all genuine exposure sits.
That last point deserves emphasis. The realistic data risk in most organisations is not the governed enterprise tool operating as designed. It is the personal account nobody knows about, holding six months of client material, on an email address the company does not control.
Frequently asked questions
Is our data used to train AI models?
At business and enterprise tiers the major providers generally exclude customer content from training by default; consumer tiers often do not. Confirm it against the contract for your specific tier and keep the wording on file.
Can we use AI with client confidential information?
That depends on your client contracts more than on the AI provider. Read the confidentiality and subprocessor clauses, identify clients with bespoke terms, and ask where the position is genuinely unclear.
How long are conversations kept?
It varies by provider and tier and is usually configurable at business tiers. Treat the default as an unmade decision rather than an accepted setting.
Is AI riskier than the cloud tools we already use?
Not inherently. It is another third-party processor and responds to the same due diligence. What makes AI feel different is that adoption often began informally, outside the process you would normally apply.
Get the data questions settled properly
Provider selection with the data terms checked, retention and sharing configured, and personal accounts folded into one governed setup: that is what the Implementation Package delivers, in 30 days or less at a fixed price.