Agentic AI Cost Tests Should Come Before Workflow Rollout

ROI & Cost Optimisation

28 September 2026 | By Ashley Marshall

Quick Answer: Agentic AI Cost Tests Should Come Before Workflow Rollout

UK businesses should run cost tests before rolling out agentic workflows, not after the first invoice arrives. A useful test measures cost per completed outcome, including tokens, tool calls, retries, human review and failed runs.

Cheaper tokens do not make agentic workflows cheap. The bill moves from model price to the number of steps, retries, tool calls and checks you let an agent run.

The cost risk has moved from model price to workflow design

For the last two years, many AI budget conversations have focused on model price. That made sense when teams were comparing one chatbot against another, or deciding whether a task needed a frontier model at all. Agentic AI changes the shape of the bill. A single user request may trigger planning, retrieval, tool calls, validation, retries, handoff messages and a final response. Each step may be individually affordable, but the workflow can still become expensive because there are so many steps.

This is the point Gartner has been warning about. The Register reported Gartner's forecast that agentic AI workflow costs could rise more than fivefold by 2028, even as base model prices fall. The useful lesson for UK business leaders is not that agents are unaffordable. It is that an agent without a cost test is not ready for production.

The common misconception is that falling token prices solve this problem on their own. They help, but they do not fix workflow inflation. If a customer service agent now reads six systems, checks policy, drafts a response, calls a CRM, gets blocked by permissions, retries twice and then sends the issue to a person, the cost driver is not just the price per token. It is the number of actions needed to reach a usable outcome.

What this means in practice is simple. Before rollout, define the business event the agent is meant to complete, then measure the full cost per successful completion. That metric is more useful than monthly AI spend because it tells you whether the workflow still makes financial sense when usage rises.

Inference is now the recurring cost centre

Training gets attention because it looks large and technical. Inference is often the quieter risk because it starts small, then compounds with every user, every workflow and every retry. Spheron's 2026 inference economics analysis states that industry analysts estimate 55-80 percent of enterprise AI GPU spend now goes to inference, and gives a production example where a 70B model serving 500 million tokens per day reaches about $950 per day in compute, or roughly $347,000 per year before overheads. Those figures will not match every UK SME, but the pattern matters.

The same article explains that a 70B deployment was reduced from $39,000 to $16,000 per month by changing model, runtime, infrastructure and FinOps choices. That is not a marginal saving. It is the difference between a workflow that can scale and one that gets quietly capped. Spheron's breakdown is infrastructure-heavy, but the management principle applies just as strongly to API-based teams: you need cost visibility before adoption becomes normal behaviour.

For a UK business buying SaaS, this can feel remote. You may never rent a GPU directly. Yet the economics still reach you through usage tiers, credits, fair use caps, overage invoices and revised supplier pricing. When vendors discover that agentic workflows cost more to serve than first expected, customers eventually see that in packaging. A flat seat that was acceptable for chat may be replaced by credits when workflows become tool-heavy.

The practical response is to stop treating pilot usage as representative. A small pilot often has friendly users, simple examples and careful operators. Production brings messy documents, edge cases, duplicated requests and impatient staff who click again when nothing appears to happen. Your cost test should deliberately include those behaviours.

A useful cost test measures failed work as well as successful work

The easiest way to understate agentic AI cost is to count only successful runs. That is rarely how business systems behave. Agents hit missing permissions, partial data, stale CRM records, rate limits, ambiguous instructions and policy conflicts. Some of those failures are healthy. You want an agent to stop when it lacks authority. But the cost of stopping still belongs in the business case.

CloudZero's 2026 AI cost optimisation guide makes the point in financial terms: AI workloads can cost from roughly $1 to $98 per GPU hour and from $0.075 to $15 per million tokens, depending on model, instance type and routing. It also notes that forgotten high-end GPU resources or bloated context windows can create significant avoidable spend. CloudZero's analysis is aimed at cloud and platform teams, but the lesson for operations leaders is broader: the waste is often in ownership, attribution and workflow habits.

A strong cost test separates four numbers. First, the median cost of a normal successful run. Second, the cost of the worst 10 percent of successful runs, because those show where complexity hides. Third, the cost of failed or escalated runs. Fourth, the human time needed to check, repair or complete the work. If the failed runs are expensive, the workflow may still be worth doing, but it needs a different design.

What this means in practice is that cost testing should sit beside quality testing. Do not only ask whether the agent gave the right answer. Ask how many steps it took, how many tokens it used, how many systems it touched, how often it retried, how often a person intervened and whether the completed outcome justified the full cost.

The business case needs unit economics, not an AI budget line

An AI budget line tells finance what was spent. Unit economics tells the business whether the spend was useful. That distinction becomes important when agentic workflows move from experiments to everyday operations. A workflow that costs 12p per completed supplier chase may be excellent if it prevents missed renewals or frees an administrator. A workflow that costs £3.80 per low-value enquiry may be poor even if the monthly invoice still looks manageable.

Telefonica Tech's guidance on AI investment cases is useful here because it separates model costs, data costs, cloud platform costs and operational costs. Its article argues that AI projects need to account for training, inference, data preparation, storage, integrations, monitoring and operational support. That is exactly the right framing for agentic work, because the agent is rarely the whole system. It is usually a coordinator sitting on top of data, tools, permissions and people.

For UK SMEs, this does not need to become a heavyweight finance exercise. Start with a simple table. List the workflow, the trigger, the expected monthly volume, the average cost per run, the escalation rate, the staff minutes saved, the error reduction expected and the commercial value protected or created. Then add a stress scenario where volume doubles, failure rates rise or a model upgrade changes token consumption.

The counterargument is that this slows innovation. In reality, it speeds up the right kind of innovation. Teams can move faster when they know which workflows are cheap enough to run freely, which need approval gates and which should stay as assisted human tasks. Cost tests are not bureaucracy when they help leaders avoid backing the wrong workflow at scale.

Routing and limits should be designed before access expands

Agentic workflows need a routing policy before broad rollout. Not every task needs the same model, context window, autonomy level or retry budget. A simple classification task may be handled by a cheaper model. A regulated customer response may need a stronger model, stricter retrieval, human review and a lower retry limit. The mistake is letting one default configuration handle everything because it worked in the pilot.

There are three controls worth putting in place early. The first is model routing by task type, so routine work does not automatically use the most expensive reasoning model. The second is a token and tool-call budget for each workflow, so a stuck agent stops and escalates instead of wandering through systems. The third is cost attribution by workflow, team and customer segment. Without attribution, leaders only see total spend after the fact.

This is also where governance and cost meet. If an agent has permission to call a CRM, send email, query finance records and update a case management system, cost is not the only risk. The workflow needs access controls, audit logs and a clear owner. Internal links between governance and economics matter here: a gateway policy can be both a safety control and a spend control, while cost allocation tags turn agent usage into evidence.

What this means in practice is that a cost test should produce operating rules, not just a pass or fail verdict. Define which model can be used, how many attempts are allowed, when a human is pulled in, which systems the agent may touch and what metric proves the workflow is still worth running. Then review those rules monthly while usage is growing.

The rollout decision should include a stop rule

A production agent should have a stop rule before it has a launch date. That may sound negative, but it is a sign of mature operating design. A stop rule says what level of cost, failure, escalation or customer impact will pause the workflow for review. It prevents slow drift, where an agent starts as a tidy efficiency project and gradually becomes a costly exception machine that nobody wants to own.

The stop rule should be tied to value, not just spend. For example, a workflow might pause if cost per completed case rises above a threshold for two weeks, if more than 15 percent of runs need manual repair, if average human review time exceeds the time the workflow was meant to save, or if a supplier changes pricing. Those triggers make the workflow accountable without requiring constant executive attention.

This is where the board-level conversation should land. The question is not whether agentic AI is good or bad. The question is which workflows have the economics, controls and ownership to run repeatedly. A UK business does not need to wait for perfect certainty, but it does need enough evidence to avoid scaling a workflow whose cost curve is already visible.

The businesses that do this well will still adopt agents. They will just adopt them with clearer boundaries. They will test the messy cases, tag the spend, compare completed outcomes, set limits and keep humans in the places where judgement still protects margin. That is the difference between AI enthusiasm and operational confidence.

Frequently Asked Questions

What is an agentic AI cost test?

It is a pre-rollout test that measures the full cost of an agent completing a business outcome, including tokens, tool calls, retries, failed runs, human review and escalation time.

Why are agentic workflows more expensive than chatbots?

They usually perform more steps. A chatbot may answer once, while an agent may plan, retrieve data, call tools, validate results, retry and hand over to a person.

Should small businesses avoid agentic AI because of cost?

No. They should start with bounded workflows where the value per completed task is clear, then measure cost before expanding access.

What metric should finance ask for?

Ask for cost per successful completed outcome, plus the cost and rate of failed or escalated runs. Monthly AI spend alone is too blunt.

How often should AI workflow costs be reviewed?

Weekly during pilot and early rollout, then monthly once usage is stable. Review immediately if pricing, model routing or workflow volume changes.

Do token prices still matter?

Yes, but they are only one part of the economics. Workflow length, context size, retry behaviour, tool access and human repair time can matter more.

What is a sensible stop rule for an AI agent?

A stop rule could pause the workflow if cost per completed task rises above target, if manual repair exceeds expected savings, or if escalation rates breach an agreed threshold.

Who should own agentic AI cost control?

Ownership should be shared between the workflow owner, finance and technical lead. The workflow owner defines value, finance checks economics, and the technical lead controls routing and limits.