AI Live Testing Evidence Should Come Before Financial Services Rollout

AI Trust & Governance

1 October 2026 | By Ashley Marshall

Quick Answer: AI Live Testing Evidence Should Come Before Financial Services Rollout

The FCA's AI Live Testing and Supercharged Sandbox programmes point towards a more demanding standard for AI evidence. UK firms should require scenario tests, consumer outcome measures, resilience results and clear accountability before moving an AI system into live financial services workflows.

A successful AI pilot is no longer persuasive on its own. Financial services buyers need evidence that the system behaves safely under realistic pressure, with real controls and named owners.

The standard is shifting from promising use case to defensible evidence

Financial services leaders have spent years hearing that AI adoption is inevitable. The more useful question now is what evidence should be required before a system touches a customer journey, payment decision, fraud alert or regulated communication. A polished demonstration proves that a model can produce an impressive result once. It does not prove that the surrounding service will remain safe, available, explainable and fair when inputs are messy, demand spikes or a supplier changes its model.

The direction of travel is visible in the Financial Services AI Adoption Plan, published by HM Treasury in July 2026. It reports that around 75% of firms surveyed by the FCA and Bank of England were already using AI, compared with 21% of firms in financial services and real estate in the wider DSIT adoption survey and 16% across the economy. Adoption is therefore not the main differentiator. The quality of testing, governance and operational evidence is becoming the differentiator.

The FCA's AI Lab reinforces that point. AI Live Testing lets firms work with the regulator on real use cases, while the Supercharged Sandbox supports controlled experimentation with compute, tools, data and expert input. These services are not a compliance certificate, and participation does not transfer accountability to the regulator. They do, however, show what mature adoption looks like: a use case is made testable, risks are surfaced early and decisions are backed by evidence rather than confidence.

For a UK buyer, this changes the approval conversation. Ask what conditions were tested, which failures were observed, what consumer outcome measures were used and who accepted the remaining risk. If the supplier can only show a benchmark score and a scripted demo, the evidence is not ready for a live financial workflow.

Testing must follow the whole service, not just the model

Many AI assurance packs focus on the model in isolation. They compare accuracy, latency or hallucination rates against a test set, then treat a good score as permission to deploy. That is too narrow for financial services. A customer experiences a service made from data pipelines, prompts, retrieval sources, policy rules, user interfaces, human reviewers, third-party APIs and escalation routes. A failure in any one of those parts can turn an acceptable model output into a poor customer outcome.

The FCA describes its Supercharged Sandbox as a place to develop and test AI use cases using GPU-enabled infrastructure, enterprise tooling, synthetic datasets and expert support. Its second cohort selected 21 organisations from 199 applications, a 51% increase on the first cohort. Those numbers matter because they show both appetite and scarcity. Most firms will not get direct access to a regulatory programme, so they need to reproduce the discipline internally.

Start with an end-to-end scenario catalogue. For an AI collections assistant, test vulnerable customer language, conflicting account records, unsupported requests, prompt injection, service outages and cases that must move immediately to a human. For an underwriting assistant, test incomplete evidence, protected-characteristic proxies, drift between customer groups and disagreement between the model and established policy. Record expected behaviour before running the test, then retain the input, output, system version, reviewer decision and remediation.

What this means in practice is that procurement cannot own the evidence pack alone. The product owner should define the intended outcome, compliance should define non-negotiable boundaries, operations should define recoverability, security should test abuse routes and customer teams should judge whether explanations are usable. A model score belongs inside that pack, but it cannot substitute for the pack.

Consumer outcome evidence needs to be designed before deployment

AI in financial services creates a particular evidence problem because a technically correct answer can still produce a poor outcome. A recommendation may be factually defensible but badly timed, difficult to understand or inappropriate for a vulnerable customer. A fraud control may reduce losses while generating so many false positives that customers lose access to essential funds. These are service outcomes, not abstract model properties.

The FCA's Mills Review was informed by research including a survey of more than 5,000 UK retail financial services consumers. HM Treasury's adoption plan highlights the tension clearly: regulated advice currently reaches only around 9% of UK adults, one in five adults are open to AI making decisions for them, and around 26% trust general-purpose tools such as ChatGPT, Claude or Gemini for financial advice. Demand is developing faster than many firms' controlled routes to serve it.

That does not mean regulated firms should avoid AI until every uncertainty disappears. It means they should define measurable customer protections before a pilot becomes a product. Useful measures include the rate of incorrect escalation, the proportion of explanations understood by users, outcome differences between relevant customer groups, successful human handoffs, complaint themes and the time taken to reverse an incorrect decision. Teams should also identify red-line events, such as fabricated eligibility criteria or the failure to recognise financial vulnerability, that automatically stop or restrict the service.

In practice, create a small outcome scorecard and review it at a fixed cadence. Do not bury it inside a technical dashboard. Give a named senior owner the job of deciding whether results justify expansion, restriction or withdrawal. Where agentic payments or other action-taking systems are involved, connect the scorecard to a clear liability map, as explained in our guide to agentic payment liability. Evidence only changes behaviour when somebody has the authority to act on it.

Resilience and concentration risk belong in the test plan

AI assurance often stops once an answer has been judged accurate. Financial regulators are looking further. The Bank of England's July 2026 Financial Stability Report says rapid advances in frontier AI have increased cyber and operational resilience risks. HM Treasury's adoption plan also warns that dependence on a small number of global cloud and model providers creates concentration, data security and service continuity concerns.

These are not remote systemic issues that only large banks need to consider. A broker, lender, insurer or fintech can create a single point of failure by placing several workflows behind one model API, one vector database or one identity service. A supplier incident can then affect customer communication, fraud handling and internal decision support at the same time. A model update can also alter outputs without the buying firm changing its own code.

The test plan should therefore include supplier failure and change scenarios. Remove access to the primary model and observe whether the workflow fails safely. Delay a third-party response and check whether staff can recognise an incomplete result. Change the model version and rerun the agreed evaluation set. Test rate limits, regional outages, corrupted retrieval data and the loss of an external moderation service. Record recovery time, data loss, manual capacity and the conditions required to return to normal operation.

The common objection is that this level of testing slows innovation. In reality, it separates reversible experiments from operational commitments. A low-impact drafting assistant may justify a light control set. A system that recommends, decides or acts in a regulated journey needs deeper evidence. Proportionality is not the absence of testing. It is matching test depth to potential harm, reversibility, customer reach and dependence. That approach lets teams move quickly where failure is containable while being deliberately cautious where it is not.

Supplier due diligence should ask for reproducible results

Buyers cannot outsource accountability, but they can make suppliers provide evidence that is useful. The usual questionnaire asks whether a vendor has a responsible AI policy, security certification or human oversight. Those answers describe intentions and organisational controls. They rarely show how the purchased system behaved in conditions relevant to the buyer.

A stronger request starts with reproducibility. Ask the supplier to state the model and service versions tested, the dates, the evaluation method, the characteristics of the data, the acceptance thresholds and the failures found. Require separate results for the exact features being purchased. Evidence for a general chatbot does not establish that an autonomous claims workflow is safe. Evidence from US customer data does not automatically demonstrate suitability for a UK regulated journey.

The buying firm should retain its own acceptance suite as well. Include representative cases, known edge cases, prohibited actions and deliberately hostile inputs. Run it before go-live, after material configuration changes and when the supplier changes a model or dependency. Contract terms should define notification periods, access to relevant logs, incident support, data handling, subprocessor changes, exit assistance and the evidence required before a significant update is accepted.

What this means in practice is a staged approval. Stage one permits controlled experimentation with synthetic or properly governed data. Stage two permits limited live use with close human review and rollback. Stage three permits wider operation only after outcome, security and resilience thresholds are met. Each stage should have an owner, expiry date and explicit scope. This is more useful than a single committee decision labelled approved because it recognises that evidence grows over time and can deteriorate after a change.

The best suppliers will welcome this clarity. It gives them a stable definition of done and reduces late objections. Suppliers that resist version disclosure, test detail or incident transparency are signalling that the buyer may carry risks it cannot observe.

A practical evidence pack for the next approval meeting

Firms do not need to copy the FCA's infrastructure to adopt the discipline behind live testing. They need a compact evidence pack that makes the deployment decision inspectable. Begin with a one-page service map showing the customer or staff journey, every automated decision, the data used, external dependencies, human checkpoints and fallback route. Add a risk-tier statement explaining why the chosen level of testing is proportionate.

Next, include five evidence groups. First, capability evidence: task success, accuracy and known limitations on representative cases. Second, outcome evidence: customer understanding, group-level differences, escalation quality, complaints and red-line events. Third, control evidence: access restrictions, logging, human authority, prohibited actions and change management. Fourth, resilience evidence: outage behaviour, recovery, manual capacity and supplier concentration. Fifth, accountability evidence: named owners, review cadence, incident decisions and the conditions that trigger rollback.

Keep the pack alive. A quarterly PDF that records an old model version is not operational assurance. Link each control to an observable measure, keep test artefacts, and rerun the acceptance suite after any change that could alter behaviour. That includes prompt changes, retrieval updates, policy rules, model upgrades, new tools and altered permissions. Review near misses as seriously as confirmed harm because they reveal where controls nearly failed.

Finally, record the decision in plain English: what the system may do, what it may not do, who is watching it, what evidence supports the decision and when approval expires. The FCA's innovation programmes are useful signals, but no external sandbox can make that decision for your organisation. The point is not to produce more paperwork. It is to give leaders enough reliable evidence to expand good systems, stop weak ones and explain either choice to customers, regulators and boards.

The immediate action is simple: take the next AI proposal off the slide deck and put it through one realistic adverse scenario. If the team cannot show the expected behaviour, observed result, owner and recovery route, the proposal is not ready for live use.

Frequently Asked Questions

Does FCA AI Live Testing approval mean an AI product is compliant?

No. FCA engagement or sandbox participation is not a certification and does not transfer accountability away from the regulated firm. The firm must still assess its obligations, customer outcomes and operational risks.

Can a smaller financial services firm apply the same approach without a large test lab?

Yes. Start with a service map, a small set of representative and adverse cases, clear acceptance thresholds, named reviewers and a tested fallback. The discipline matters more than expensive infrastructure.

What should an AI supplier provide before a live pilot?

Ask for versioned test results, known limitations, data and evaluation details, security and incident controls, update notification terms, relevant logs and evidence for the exact feature being purchased.

How often should an AI system be retested?

Retest on a regular risk-based schedule and after any material change to the model, prompt, retrieval data, policy rules, connected tools, permissions or supplier dependencies.

What is a red-line event in AI testing?

It is a predefined failure serious enough to stop, restrict or roll back the service, such as fabricated eligibility rules, an unauthorised transaction or failure to recognise a vulnerable customer.

Is human oversight enough to make an AI workflow safe?

Not by itself. Reviewers need time, context, authority, training and a workable escalation route. Firms should test whether humans actually detect and correct failures under realistic workload.

Does this approach only apply to banks?

No. Insurers, lenders, brokers, advisers, fintechs and suppliers can use the same evidence model, with test depth proportionate to customer impact, reversibility and regulatory exposure.