How can I tell whether an AI tool is using my business data to train its models?

5 October 2026

How can I tell whether an AI tool is using my business data to train its models?

Do not rely on the supplier's homepage or a salesperson saying your data is secure. Verify the terms for the exact free, individual, business or enterprise plan you will use, check whether training is off by default or merely optional, and record retention, human review, subprocessors and opt-in settings. If the answer is missing, qualified or spread across conflicting documents, treat the tool as unsuitable for confidential business data until the supplier answers in writing.

Start with the exact product and plan, not the brand name

The same AI brand can apply very different data rules to a free personal account, a paid individual account, a business workspace and an API. Your first job is therefore to write down the exact product, plan and sign-in route your staff will use. A promise made for an enterprise service does not automatically cover a free chatbot account, a browser extension, a third-party integration or a custom tool built on the same model.

Look for an explicit sentence that answers: are our inputs and outputs used to train or improve the supplier's models? Inputs include prompts, documents, images, audio, connected mailbox content and information retrieved from business systems. Outputs are the answers the service generates. Then look for qualifications such as "by default", "without permission", "may be reviewed" or "unless you provide feedback". These words are not necessarily red flags, but they tell you which settings and behaviours can change the answer.

For example, OpenAI's business data page says that inputs and outputs from ChatGPT Business, Enterprise, Edu and its API platform are not used for model training by default. That is useful evidence for those named services. It is not evidence that every personal account, custom GPT, connector or external app follows identical rules. Microsoft similarly says protected Copilot prompts and responses are not used to train underlying foundation models, while Google says qualifying Workspace content is not used for generative AI model training outside the customer's domain without permission. The product boundary matters.

The Office for National Statistics reported in July 2026 that 35% of UK businesses with at least 10 employees used at least one AI technology, up from about 12% in late 2023. It also found that 55% of employees reported using AI for work or education. The gap is a warning: individual use can spread faster than a business can formally assess it. Your review must cover what staff actually use, not only what the company has purchased.

Check five places before approving the tool

Start with the privacy notice, but do not stop there. The most reliable review checks five places: the product privacy page, the terms of service, the data processing agreement, the administrator controls and the documentation for any connected feature. Search each document for "train", "improve", "retention", "human review", "feedback", "subprocessor" and "delete". Save a PDF, screenshot or dated link because terms and settings change.

First, the privacy page should explain the general treatment of prompts, files and responses. Second, the contract or terms should state which document takes priority and whether the supplier acts as your processor for personal data. Third, the data processing agreement should identify processing purposes, deletion arrangements, international transfers and subprocessors. Fourth, the admin console should show whether users can enable feedback sharing, conversation history, connectors, public sharing or third-party extensions. Fifth, each integration needs its own check because a tool may send data to another provider even when the main AI supplier does not train on it.

Do not confuse encryption with a no-training promise. Encryption at rest and in transit protects data against some forms of unauthorised access. It does not tell you whether the provider is contractually permitted to analyse the content, retain it, let reviewers inspect it or use it to improve a model. Likewise, "we do not sell your data" does not answer the training question.

Retention deserves a separate line in your decision record. Google's August 2026 Workspace Privacy Hub, for example, says Gemini in Workspace prompt and response retention can range from 90 days to indefinite depending on administrator choices, while the Gemini app offers different controls and periods. A tool can promise not to train on your data and still retain it for service delivery, safety, abuse monitoring or legal reasons. You need both answers: is it used for training, and how long does it remain available?

Understand what the major suppliers actually promise

The mainstream business products provide useful reference points. OpenAI states that it does not train on organisation data from its named business products and API by default, and that API customers may explicitly opt in to share data. It also describes encryption using AES-256 at rest and TLS 1.2 or higher in transit, plus retention controls for qualifying organisations. The practical check is whether your account is genuinely inside one of those business products and whether an owner has enabled any voluntary sharing.

Microsoft's Copilot privacy documentation says prompts and responses covered by enterprise data protection are not used to train underlying foundation models. Microsoft 365 Copilot can still access information that a signed-in user is already permitted to see through Microsoft Graph. That creates a different risk: excessive SharePoint, Teams or OneDrive permissions may cause Copilot to surface information to someone who technically has access but would not normally find it. A no-training promise does not repair poor internal permissions.

Google's Workspace Privacy Hub says qualifying Workspace interactions stay within the organisation, existing controls apply and content is not used to train generative AI models outside the domain without permission. It also states that Gemini accesses relevant Workspace content based on the user's prompt and existing permissions. Again, review sharing and access before connecting AI.

These are named examples, not blanket endorsements. Suppliers revise products, plans and language. Smaller vendors may route prompts through OpenAI, Anthropic, Google, Microsoft or another model provider while storing copies in their own database. Ask the vendor to identify every model provider and subprocessor, then confirm which contract governs each hop. The weakest link may be the application vendor rather than the model company printed on its marketing page.

Ask the supplier these questions and require written evidence

Send the supplier a short written questionnaire. Ask whether prompts, uploads, responses, metadata, feedback and support conversations are used to train, fine-tune, evaluate or improve any model. Ask whether the default differs by plan, region or feature. Ask who can access content, where it is processed and stored, how long it is retained, whether you can delete it, which subprocessors receive it and what happens when an employee leaves.

Then ask about exceptions. Many services exclude normal business content from training but may use information submitted through thumbs-up feedback, bug reports, support tickets, beta features or voluntary data-sharing programmes. A staff member can accidentally change the data treatment by clicking an opt-in control. Your approval should specify which settings must remain off and who is allowed to change them.

Request the privacy notice, current data processing agreement, subprocessor list, security documentation and a screenshot or demonstration of the relevant admin controls. Certifications such as ISO 27001 and SOC 2 can support a security review, but they do not answer the model-training question on their own. The decisive evidence is clear contractual language plus settings you can inspect and control.

If a salesperson answers "we are GDPR compliant", ask again. Compliance is not a feature switch and the phrase does not explain lawful basis, controller and processor roles, retention, international transfers or training. If the supplier will not answer in writing, restrict the trial to invented or fully anonymised information. Do not upload customer files, employee records, contracts, credentials, source code or commercially sensitive plans just to see whether the tool works.

A simple decision can be green, amber or red. Green means clear no-training language for your plan, acceptable retention, a signed processing agreement and controlled access. Amber means the tool may be useful with synthetic or low-sensitivity data while questions are resolved. Red means unclear terms, training enabled by default, no business contract, uncontrolled sharing or no workable deletion process.

Your UK GDPR responsibilities do not disappear when training is off

A no-training commitment is valuable, but it does not make every use lawful or sensible. If prompts contain personal data, your business still needs a purpose, a lawful basis, appropriate transparency, data minimisation, security and a suitable processor contract. You must also consider accuracy, access rights, deletion and whether people could be significantly affected by an AI-assisted decision.

The Information Commissioner's Office guidance on AI and data protection centres accountability, governance, transparency, lawfulness, accuracy and fairness. The guidance is under review following the Data (Use and Access) Act, so check the current version when making a material decision. For a high-risk use, document whether a data protection impact assessment is required and involve whoever is responsible for data protection in your business.

The UK Government's Cyber Security Breaches Survey 2025/2026 found that 31% of businesses were using, adopting or actively considering AI, but only 24% of that group had cyber security practices or processes in place to manage AI risks. The same survey found that 43% of businesses had identified a cyber breach or attack in the previous 12 months. Those figures do not prove an AI tool caused a breach. They do show why a vague supplier promise should not replace normal risk management.

Minimise before you upload. Remove names, email addresses, account numbers, health details and unnecessary commercial information. Use a short extract rather than a full mailbox or drive. Give connectors read-only, narrow access where possible. Keep human review for outputs that affect customers, staff, money, legal rights or safety. Record the tool, owner, purpose, data categories, approved users, settings, review date and exit plan in your AI register.

A practical 30-minute approval check

Spend the first five minutes defining the use case. Name the users, the exact plan, the information they want to enter and the business result expected. If nobody can describe the data clearly, stop. "General productivity" is not enough because summarising a public report is very different from analysing employee grievances or customer case files.

Use the next ten minutes to find and save the supplier's training statement, retention terms, processing agreement and subprocessor list. Mark each claim as confirmed, unclear or unacceptable. Check whether the wording covers prompts, uploads and outputs, and whether feedback or optional sharing creates an exception. If you are using an app built on another company's model, repeat the check for the app provider and underlying model provider.

Use another ten minutes in the admin console. Turn off optional data sharing, public links and unneeded connectors. Require company-owned accounts, multi-factor authentication and an accountable workspace owner. Check who can add integrations and whether accounts can be removed promptly when staff leave. Test deletion with a harmless conversation or sample file rather than assuming the button works as expected.

Use the final five minutes to record the decision and tell staff the boundary in plain English. For example: "Approved for drafting from public or internal low-sensitivity material. Do not enter customer personal data, contracts, passwords, payment information or HR records. Use only the company workspace. Report accidental uploads immediately." Set a review date within three to six months, or sooner if the supplier changes terms, launches a new connector or suffers an incident.

If the evidence is incomplete, run a low-risk trial with invented data. You do not need to reject every imperfect tool, but you do need to match the data to the evidence. That is the honest answer: trust is not created by a logo or a privacy slogan. It comes from specific terms, controlled settings, limited access and a record showing why the decision was reasonable.

When this does not apply

This lightweight check does not apply where the proposed use could create a high risk to people's rights, reveal special category data, make or strongly influence employment decisions, provide regulated advice, process children's information or expose large collections of client records. It also does not cover training your own model, fine-tuning a model on customer data or building an agent with broad access to email, finance and operational systems. Those uses need a deeper technical, contractual and data protection assessment.

Do not assume anonymisation is easy. Removing a name may leave enough detail to identify a person from their job title, case history, location or combination of facts. Pseudonymised data is still personal data if your organisation can reconnect it to an individual. When genuine anonymisation cannot be assured, treat the material as personal data.

If you need the tool only for public research, generic brainstorming or drafting from non-confidential notes, a narrow approval may be enough. If you need it to work with sensitive business information, pay for the managed product and controls that fit the risk, or choose a different approach. Sometimes the right answer is a conventional search, template, rule-based automation or local processing rather than a general-purpose chatbot.

If you want help assessing a specific product, prepare the plan name, intended use, sample data categories and links to its current terms before speaking to an adviser. That makes the review faster and more useful. A good adviser should tell you what is confirmed, what remains uncertain and what would make the use unacceptable. No pitch, no pressure, just a documented decision you can defend.

Is This Right For You?

This check is right for any UK business considering an AI tool that may receive customer emails, contracts, meeting notes, financial information, staff records, source code or internal plans. It is especially useful when staff already use a consumer account and you are deciding whether to ban it, replace it with a managed business workspace or approve a narrow use case.

It is not a substitute for legal or data protection advice where you process special category data, large volumes of personal data, children's data, regulated professional information or information that could cause serious harm if exposed. In those cases, involve your data protection lead or adviser and consider a data protection impact assessment before deployment.

Frequently Asked Questions

Does turning off chat history stop an AI company training on my data?

Not necessarily. History, retention and model training are separate controls. Check the supplier's wording for your exact plan and confirm whether turning history off changes retention, human review or training treatment.

Is ChatGPT Business data used to train OpenAI models?

OpenAI states that ChatGPT Business inputs and outputs are not used to train or improve its models by default. Confirm that staff are using the managed Business workspace, review voluntary sharing settings and check the current terms before approval.

Does Microsoft 365 Copilot train on company emails and files?

Microsoft states that prompts, responses and Microsoft Graph data covered by enterprise data protection are not used to train foundation models. Copilot can still retrieve content a user is permitted to access, so review Microsoft 365 permissions as well as training terms.

Does Gemini for Google Workspace use business data for training?

Google states that qualifying Workspace content is not used to train generative AI models outside the customer's domain without permission. Check that your edition qualifies and review admin settings, conversation retention and connected services.

Can I safely test an AI tool before the supplier answers my questions?

Yes, but use invented, public or genuinely anonymised information only. Do not upload customer files, staff records, contracts, credentials, financial data or confidential plans during an unapproved trial.

What should I do if the terms say data may be used to improve services?

Ask whether that includes model training, evaluation, human review or only operational diagnostics. Require a written answer, check for an opt-out and restrict the tool to low-sensitivity data until the scope is clear.

Do I need a data protection impact assessment for every AI tool?

No. A DPIA is required where processing is likely to result in a high risk to people's rights and freedoms. Use the ICO screening criteria and take advice for sensitive data, monitoring, profiling or consequential decisions.