AI Agent Security Evidence Should Be In Supplier Contracts Before UK Teams Buy
Agentic Business Design
12 September 2026 | By Ashley Marshall
Quick Answer: AI Agent Security Evidence Should Be In Supplier Contracts Before UK Teams Buy
UK businesses buying agentic AI should make security evidence a contract requirement before a pilot starts. Ask for documented sandbox boundaries, logging, human oversight, incident response duties and change notices, then attach those artefacts to acceptance gates.
The next AI procurement risk is not whether an agent looks clever in a demo. It is whether the supplier can prove how it will be constrained, monitored and stopped when it behaves unexpectedly.
The buying question has moved from capability to controllability
For the last year, most AI agent conversations have started with capability. Can the tool read the inbox, open the CRM, draft a response, update the finance record, reconcile the ticket and chase the supplier? That is the wrong first question for any UK organisation planning to put agents near real business systems. The better question is whether the buyer can see and control what the agent is allowed to do.
The National Cyber Security Centre has now made that practical point very clearly. Its August 2026 guidance on managing the cyber risk of agentic AI says organisations should consider how these systems are deployed, constrained, observed and responded to, especially where agents may act with significant autonomy. That is procurement language as much as security language. If a supplier cannot explain the boundary of the agent, the logging model, the approval model and the stop procedure, the buyer is not ready to let that agent touch live work.
What this means in practice is simple: supplier contracts should ask for evidence, not reassurance. A demo can show a happy path. A contract pack should show the operating boundary, the failure cases, the testing method and the named owner for each control. This is not about slowing adoption. It is about avoiding the expensive pattern where a pilot wins internal enthusiasm before anyone has checked whether it can be governed in production.
The useful shift is to treat controllability as a buying criterion. Before a UK team chooses an agent platform, it should ask what the agent can access, what it cannot access, what events are logged, which actions need approval, how unexpected behaviour is detected and who can pull the plug. Those answers belong in the commercial process, not in a security review after the purchase order has gone through.
Ask suppliers for their sandbox and access boundary evidence
The first evidence pack should describe the agent's operating environment. NCSC advises that AI agents should run within a sandboxed environment that controls what resources can and cannot be communicated with, locally and over a network. That matters because an agent's risk is not limited to the model. It includes the browser session, connector permissions, API tokens, file system access, email scope, SaaS roles and any route from one system into another.
A supplier should therefore provide a clear access boundary before implementation starts. Buyers should expect a system diagram, not just a paragraph in a security questionnaire. It should show where the model runs, where the orchestration layer runs, which tools the agent can call, which networks it can reach, which identities it uses, and how access is segregated between test and production. If the supplier says the customer controls all permissions, that is not enough. The buyer still needs to know how the product enforces least privilege and prevents accidental expansion.
This is where many agent pilots go soft. Teams grant wide access during a proof of concept because it is faster, then forget that the permissions were temporary. The contract should make temporary access explicit. It should require a permission review before production, a connector inventory, and a documented route for removing access when the pilot ends. For Microsoft 365, Google Workspace, Salesforce, HubSpot, Xero or sector systems, that inventory should list the exact scopes and roles being used.
The counterargument is that tight sandboxing makes agents less useful. Sometimes that is true. An agent that can only work with a narrow dataset may complete fewer tasks. But that is a design trade-off, not a reason to skip the control. Start with the smallest useful boundary, measure what fails, then expand access deliberately. A supplier worth buying from should be able to support that staged model.
Make logs, attribution and monitoring part of acceptance
Agentic AI changes the evidence requirement because actions can happen across several systems and several steps. A human user might approve a refund, update a CRM field and email a customer in three visible moments. An agent may interpret a goal, call tools, retry a failed step, summarise a result and trigger another action. If the buyer cannot reconstruct that chain afterwards, it has created an operational blind spot.
NCSC's guidance calls for observability, audit and monitoring of agentic AI activity as part of security operations. It also says activity should be easy to attribute. In procurement terms, that means the supplier should be able to show what gets logged, how long logs are retained, how logs can be exported, and whether each action can be tied back to a user, workflow, agent identity, tool call and approval event. For regulated or customer-facing processes, vague platform analytics are not enough.
Acceptance criteria should include a log review. Before signing off a pilot, ask the supplier to run a controlled scenario and provide the evidence trail. Can your team see the original instruction, the data accessed, the tool calls made, the approvals requested, the actions completed, the failed attempts and the final output? Can security operations or the process owner receive alerts when the agent touches a sensitive record, exceeds a spend threshold or attempts an out-of-scope action?
This also matters for ROI. Without logs, you cannot tell whether an agent is genuinely saving time or merely moving effort into exception handling. The same trail that helps security investigate an incident can help operations spot rework, failed automations and low-value tasks. That is why logging should not be treated as a compliance afterthought. It is the management evidence for whether the agent is working.
Incident duties need to be explicit before something goes wrong
Most AI supplier terms are still written as if the main risk is service availability or data processing. Agentic systems add another question: what happens when the system takes an unintended action? NCSC's 4 August 2026 statement on frontier AI incidents warned that relying on detection after the fact will not be enough and called for strong safeguards, real-time oversight and clear plans for responding when the unexpected happens.
That wording should change how UK buyers negotiate agent contracts. The supplier should commit to incident notification duties, evidence preservation and support for root cause analysis. The contract should define what counts as an AI agent incident: unsanctioned tool use, access outside scope, unexpected data disclosure, unauthorised external communication, incorrect high-impact action, control bypass, repeated unsafe recommendation, or any behaviour that breaches the agreed operating boundary.
For organisations with EU exposure, the EU AI Act also raises the temperature. Article 73 requires providers of high-risk AI systems placed on the Union market to report serious incidents. Even where a UK buyer is not directly in scope for a specific high-risk system, the direction of travel is clear: incident reporting, logs and post-market monitoring are becoming normal governance expectations. UK firms should not wait until a regulator forces the point. They should ask suppliers now how incidents will be detected, reported and evidenced.
What this means in practice is a short contract schedule. It should state reporting timelines, contacts, minimum evidence, cooperation duties, customer notification support and the right to pause or disable the agent. It should also require the supplier to disclose material model, policy or orchestration changes that could alter behaviour. If the supplier cannot agree to basic operational incident duties, the buyer should treat that as a production readiness gap.
Human oversight has to be designed, named and tested
Human oversight is often promised and rarely designed. A supplier says a person remains in control because the agent produces suggestions or asks for approval. The real question is whether the approval is meaningful. Who receives it? What information do they see? Can they inspect the evidence? Is approval guaranteed before the action, or is the person merely notified after the event? What happens if the approver is unavailable?
NCSC describes different oversight models, including human-in-the-loop, human-on-the-loop and human-out-of-the-loop. For higher-risk scenarios, it recommends human oversight alongside technically enforced controls. Buyers should turn that into a workflow design requirement. Each agent action should be mapped to an oversight level. Low-risk drafting may need review after generation. Customer refunds, supplier changes, payroll actions, legal commitments or record deletion should need stronger gates.
Named ownership matters as much as the gate itself. An agent used by sales operations may be configured by IT, owned by revenue operations, monitored by security and relied on by customer success. If nobody owns the decision rights, failures become everyone's problem and nobody's job. The contract and internal launch pack should name the business owner, technical owner, security contact and escalation path. That ownership should survive supplier onboarding and remain visible in the runbook.
The misconception is that humans in the loop remove the need for technical controls. They do not. People approve bad actions when the approval screen lacks context, when alerts are too frequent, or when the process trains them to click through. A good supplier should help test the oversight design with realistic cases, including ambiguous requests, missing data, permission failures and attempts to cross a boundary. If the human gate only works in a clean demo, it is not a control.
Turn supplier evidence into a production gate
The practical answer is not a longer questionnaire. It is a production gate. Before an agent moves from pilot to live use, the supplier and buyer should assemble a small evidence pack that proves the agreed controls exist. That pack should include the access boundary, connector inventory, logging sample, oversight map, incident procedure, emergency shutdown process, change notice commitment and test results for the most important failure cases.
This aligns with the UK's wider direction on AI security. The Government's AI Cyber Security Code of Practice sets out baseline cyber security principles for organisations that develop and deploy AI systems, and GOV.UK notes that DSIT and NCSC intend the code and implementation guide to form the basis for a new global ETSI standard. Buyers do not need to wait for every standard to settle before improving procurement. They can use the code's direction of travel to make evidence normal in supplier selection.
A useful production gate is specific and short. Ask five questions. What can the agent access? How do we know what it did? Which actions require approval? How do we stop it quickly? What must the supplier tell us when the system changes or behaves unexpectedly? If the answer to any question is verbal only, the gate has not been met. Evidence can be a diagram, configuration export, policy document, log sample, test result or contract clause.
This approach also keeps the conversation commercially fair. Suppliers are not being asked to reveal proprietary model internals. They are being asked to prove operational controls around a tool that may touch customer data, money, legal commitments or core records. Good suppliers should welcome that because it separates serious production platforms from impressive demos. For UK leaders, the lesson is straightforward: buy agents like operational systems, not like clever chatbots.
Frequently Asked Questions
What evidence should we ask an AI agent supplier for before a pilot?
Ask for the access boundary, connector scopes, sandbox design, logging model, approval workflow, incident procedure, shutdown process and change notice policy. For higher-risk workflows, ask for a sample evidence trail from a test run.
Is this only relevant to regulated sectors?
No. Regulated firms have a stronger evidence burden, but any UK business letting agents touch customer data, finance records, HR information, supplier portals or production systems needs proof that the agent can be controlled.
Should suppliers disclose their model internals?
Usually not. The more practical requirement is operational evidence: what the agent can access, what it can do, how activity is logged, which actions need approval and how the system can be stopped.
How does NCSC guidance affect AI agent procurement?
NCSC's guidance makes clear that agentic AI needs safeguards, sandboxing, oversight, observability and response planning. Buyers can turn those expectations into contract schedules and production acceptance criteria.
What is the biggest mistake in AI agent pilots?
The common mistake is granting broad temporary access to make the demo work, then failing to redesign permissions before production. Temporary access should be documented, reviewed and removed unless it is deliberately approved.
Do human approval steps remove the need for technical controls?
No. Human oversight works only when the approver has enough context and the system technically enforces the gate. Approval screens, logs, alerts and access boundaries need to work together.
How should we define an AI agent incident?
Define it around the agreed boundary: unsanctioned tool use, out-of-scope access, unauthorised communication, incorrect high-impact action, data exposure, control bypass or behaviour that breaches the approved workflow.
What should a production gate include?
A production gate should include evidence of access controls, logs, oversight, incident response, emergency shutdown, supplier change notices and test results for important failure scenarios.