AI Data Provenance Registers Should Come Before Model Expansion
AI Trust & Governance
6 September 2026 | By Ashley Marshall
Quick Answer: AI Data Provenance Registers Should Come Before Model Expansion
UK firms should treat data provenance as a live operating register, not a one-off compliance note. Before expanding models, adding agents or connecting new datasets, leaders need a current record of source, permission, ownership, quality, transfer history and usage constraints.
The next AI governance gap is not another policy. It is knowing exactly where the data came from, who was allowed to use it, and whether that evidence survives the next model change.
Data provenance is becoming an operating control
Most AI governance programmes still begin in the wrong place. They start with a policy, an acceptable use document, or a risk committee, then ask delivery teams to explain data decisions after the system is already live. That approach is too slow for the way AI systems now change. Retrieval stores are refreshed, copilots gain connectors, agents are given new tools, and suppliers update models without waiting for the next quarterly governance meeting.
The stronger starting point is a data provenance register: a live record of where each dataset, document store, prompt corpus, fine-tuning sample, synthetic dataset and externally supplied knowledge source came from, who controls it, what permissions apply, how it was transformed, and which AI systems now rely on it. This is not just a legal artefact. It is the operational map that lets a firm answer basic questions during procurement, incident response, model upgrade review and customer due diligence.
Recent UK policy signals point in the same direction. In July 2026, the Government's call for evidence on data regulation in the age of AI said AI creates more demanding requirements for personal and non-personal data, including clearer provenance and permissions, metadata usable by both people and machines, and governance across the AI lifecycle. That is a practical sentence for boards. It says the control problem is no longer limited to privacy notices or supplier questionnaires. It is about evidence that can travel with the data.
What this means in practice is simple: do not approve model expansion unless the data register is current enough to explain the system. If a team cannot show source, lawful basis or licence, owner, quality checks, transformation history, retention period and downstream AI dependencies, it is not ready for wider deployment. A model without a provenance register is a dependency you cannot inspect.
The evidence burden is moving from promise to record
There is a useful misconception to clear up. A provenance register is not the same as a data catalogue. A catalogue usually helps people find datasets. A provenance register helps decision makers defend the use of those datasets in an AI system. It should show who created the data, how it was collected, who changed it, whether it includes personal or sensitive attributes, what restrictions apply, and where it is used.
The UK's AI Management Essentials tool is explicit about this direction. It asks whether an organisation maintains a complete and up-to-date record of the AI systems it develops and uses. It describes an AI system record as an inventory of documentation, assets and resources that may include technical documentation, impact and risk assessments, AI model analyses and data records. It then asks whether organisations request and receive documentation, assets and resources from third-party providers for that record.
That matters because supplier-heavy AI stacks are now normal. A business may use Microsoft 365 Copilot, a specialist legal assistant, a customer support agent, an analytics tool with embedded generative AI and a custom retrieval assistant. Each one may touch different data sources. Each one may also depend on suppliers whose own records are incomplete or unavailable to customers. If the buyer does not keep a local register, it has no joined-up view of the evidence position.
The practical control is to make provenance part of intake, not audit clean-up. Every new AI system entry should include a data tab with source categories, permissions, processing purpose, sensitivity rating, quality checks, known gaps, retention rules, supplier evidence received, and named business owner. For internally trained or fine-tuned systems, add collection process, labelling process, transformation steps, exclusion rules and a review date. For retrieval systems, include each connected repository and the access model used. This turns governance from vague comfort into a repeatable release gate.
Regulators are asking for practical certainty, not theatre
The UK's current direction is not to freeze AI adoption until every uncertainty disappears. It is to create enough practical certainty for responsible adoption. That distinction is important. The ICO's June 2026 statement on the advisory AI Growth Lab says regulators will give innovators and adopters clear, practical information on how existing regulations apply to novel AI applications. The first focus is LawTech, legal services and conveyancing, where data protection, professional duties and customer trust meet quickly.
The ICO's July 2026 sandbox blog adds another useful lesson. It says a data protection statutory regulatory sandbox is feasible, but only if individual rights are protected through alternative but equivalent accountability and governance mechanisms. That phrase is doing real work. It implies that when firms push into new AI uses, they still need demonstrable mechanisms for accountability. A provenance register is one of those mechanisms because it lets teams explain what data is in scope, what is excluded, and how people can challenge or investigate the system later.
This also answers the common counterargument: "We already have a DPIA and a supplier contract, why add another register?" The answer is that those documents are necessary, but they are often static. A DPIA may describe a system at approval time. A supplier contract may set obligations at purchase time. A provenance register tracks the live state as data sources change, connectors expand and model behaviour is retested. It is the working evidence layer between policy and operations.
For UK leaders, the minimum viable version can be modest. Start with high-impact systems that affect customers, employees, regulated decisions, financial reporting, legal work or operational resilience. Add a simple red, amber and green evidence status. Red means source or permission is unclear. Amber means evidence exists but is incomplete or old. Green means the system can show current source, permission, owner, quality, transfer and downstream use. That status should be visible to the system owner, procurement, legal, security and the AI governance lead.
Sector evidence shows why provenance cannot stay generic
Generic governance language becomes weak when AI touches real operational evidence. The Food Standards Agency's June 2026 Science Council report on AI in food safety and authenticity is a useful example because it deals with assurance, inspection and documentation rather than abstract AI strategy. The report discusses AI-supported data pack generation for third-party certification and assurance, AI-powered document inspection at UK ports of entry, and AI-assisted detection in abattoirs. It says AI could make assurance faster and more consistent, but also warns that systems may embed bias or drift, generate outputs that are hard to explain or reproduce, and conceal weaknesses behind apparently robust documentation.
That last risk is directly relevant outside food. AI can make documentation look more complete than the underlying evidence really is. A supplier assurance pack can be beautifully formatted while drawing from stale records. A compliance assistant can summarise policies without knowing which ones were superseded. A sales enablement bot can quote approved claims from a document store that nobody has reviewed since the product changed. Without provenance, polished output becomes a confidence trick.
The FSA report's recommendations include promoting data quality, provenance and standards, including data and cybersecurity, to help realise fair, auditable AI. It also calls for responsible use guidance, ongoing monitoring, validation mechanisms and clear human accountability. That is exactly the operating pattern many UK firms need: provenance plus validation plus accountable ownership, not provenance as a spreadsheet that nobody revisits.
What this means in practice is that the register must include domain context. In legal services, record privilege, client confidentiality, matter boundaries and professional responsibility constraints. In HR, record protected characteristic exposure, candidate consent, retention and fairness monitoring. In finance, record source system authority, reconciliation rules and audit evidence. In customer support, record product version, policy date, complaint risk and escalation rules. The fields should be consistent enough to compare across the business, but specific enough to catch sector risks.
EU AI Act pressure will reach UK procurement
Even where a UK firm is not directly building high-risk AI for the EU market, European rules are changing buyer expectations. The European Commission's AI Act guidance describes a risk-based framework and says transparency rules come into effect in August 2026. For high-risk systems, it lists obligations including high-quality datasets, logging of activity to ensure traceability, detailed documentation, clear information to deployers, human oversight, robustness, cybersecurity and accuracy. Those are not niche concerns for lawyers. They are the language buyers will use when they compare suppliers.
UK organisations selling into regulated sectors, working with EU customers, or buying tools from global vendors should expect more questions about evidence. Can you explain the data used to train, tune or ground the system? Can you show traceability of outputs? Can you distinguish model training data, customer data, retrieval data and user prompts? Can you prove that restricted sources were not used? Can you remove or quarantine a dataset when a permission problem appears?
The mistake is to treat this as an EU compliance problem only. Procurement teams in the UK are already translating external regulation into supplier due diligence. A buyer may not quote the AI Act directly, but they will ask for model cards, data source summaries, audit logs, security attestations, testing evidence, incident processes and contractual commitments. A provenance register helps the seller respond quickly and helps the buyer decide whether the supplier's claims are credible.
For boards, the commercial point is sharper than the regulatory one. Provenance evidence reduces friction in sales, investment, insurance, audits and enterprise procurement. It also reduces the cost of investigation when an AI system produces a disputed recommendation. If the business has to reconstruct source history from Slack messages, vendor portals and scattered spreadsheets after a complaint, it has already lost time and trust. A register turns future scrutiny into a prepared evidence request rather than a scramble.
Build the register before the next connector goes live
The sensible implementation route is deliberately unglamorous. Choose a small number of mandatory fields, apply them consistently, and connect the register to approval workflows. A useful first version includes system name, business owner, technical owner, supplier, purpose, decision impact, data source, data type, personal data status, lawful basis or licence, source owner, collection method, transformation history, quality checks, access model, retention rule, downstream systems, last review date, evidence link and approval status.
Then add three operating rules. First, no new AI system moves from pilot to production without a provenance entry. Second, no new connector, dataset or supplier update is approved unless the entry is refreshed. Third, any red evidence status automatically triggers a risk review before the system is expanded. This does not need to stop teams from experimenting. It does stop undocumented data dependencies from becoming business-critical infrastructure.
Tooling can be simple at the beginning. Some firms can start in a controlled spreadsheet or governance workspace. Others will need integration with ServiceNow, Jira, Microsoft Purview, Collibra, Alation, OneTrust, Vanta, Drata, Azure AI Foundry or internal data catalogues. The tool matters less than the ownership model. Someone must be accountable for keeping each entry current, and someone else must be able to challenge weak evidence before deployment.
The final test is whether the register can answer a board-level question in plain English: "If this AI system gives a wrong, biased, unlawful or commercially damaging answer, can we trace the data path quickly enough to act?" If the answer is no, model expansion is premature. Better prompts, bigger context windows and smarter agents will not fix a missing evidence base. Provenance is not bureaucracy. It is the operating memory of the AI estate.
Frequently Asked Questions
What is an AI data provenance register?
It is a live record showing where the data used by an AI system came from, who controls it, what permissions apply, how it was changed, and which systems depend on it.
How is this different from a data catalogue?
A data catalogue helps people find and understand datasets. A provenance register is narrower and more evidence-led: it helps the business justify, audit and investigate AI use of data.
Do small UK businesses need this?
Yes, but they can start small. A controlled spreadsheet covering high-impact AI systems is better than waiting for a full governance platform.
What fields should we capture first?
Start with source, owner, licence or lawful basis, personal data status, collection method, transformation history, quality checks, retention, downstream systems and review date.
Should supplier AI tools be included?
Yes. If a third-party AI tool affects your operations, customers, staff or regulatory position, your organisation still needs a local evidence record for it.
Who should own the register?
Each AI system needs a named business owner, with support from data, legal, security and technology teams. Governance cannot sit with IT alone.
Is this required by UK law?
Not as a single named register in most cases. However, UK data protection, procurement, assurance and accountability expectations increasingly depend on records that show source, purpose, quality and responsibility.
When should the register be updated?
Update it whenever a new dataset, connector, supplier, model version, processing purpose or retention rule changes. At minimum, review high-impact entries twice a year.