Build an AI Agent Service Catalogue Before Teams Create Their Own

Agentic Business Design

9 October 2026 | By Ashley Marshall

Quick Answer: Build an AI Agent Service Catalogue Before Teams Create Their Own

Create a service catalogue that records every agent, its owner, purpose, data, tools, autonomy and retirement conditions. It gives teams a reuse route while giving leaders a practical control point for agent sprawl.

Your next AI agent problem may not be a bad model. It may be five teams building the same capability with different permissions, owners and failure rules.

Agent sprawl starts as reasonable local decisions

Most organisations do not set out to create an uncontrolled estate of AI agents. A sales team builds a research assistant, finance tests an invoice exception agent, customer service configures a response tool and IT creates a helpdesk workflow. Each decision can be sensible in isolation. The problem appears when nobody can answer three basic questions across the whole business: what agents exist, which capability each one provides and who is accountable for its behaviour.

Recent evidence suggests confidence is running ahead of visibility. A July 2026 survey of 700 technology professionals at large organisations in five countries, including the UK, found that 77% were confident they had a complete inventory of every agent, MCP server and large language model in their environment. Only 44% used active discovery tooling to verify it. The Harness State of Agent DLC findings also reported that roughly seven in eight respondents had experienced at least one tangible agent-related issue during the year. These figures come from committed adopters rather than all businesses, but that makes the warning more useful: deployment maturity does not automatically produce estate clarity.

A service catalogue closes a different gap from a security register. It records the business service as well as the technical object. For each agent, it should show the user need, service owner, intended users, approved data, available tools, decision rights, human checkpoints, support route, performance evidence and planned review date. It should also say whether another approved agent already provides the same capability. What this means in practice is that a team proposing a new research agent must first search the catalogue, assess reuse and explain any genuine difference. Discovery still matters, as outlined in our shadow AI discovery sprint, but the catalogue turns a one-off discovery exercise into an operating habit.

The catalogue must describe a service, not just a piece of software

A spreadsheet containing an agent name, vendor and URL is an asset list. It is useful, but it will not help a business decide whether to reuse, change or retire an agent. A service catalogue starts with the outcome. It describes what the agent is allowed to do for whom, the evidence that it works and the operating conditions that make its use acceptable.

The distinction matters because agentic systems combine capabilities at runtime. The UK Government's AI Insights guidance on agentic AI, updated in August 2026, explains that tools and functions are registered as capabilities and described in natural language so a model can select a course of action. It also warns that unclear descriptions can lead the system to select the wrong function. The catalogue therefore needs more than a label such as customer agent. It needs a bounded purpose such as drafts replies to delivery-status enquiries using order and courier data, but cannot issue refunds or change addresses. That description is useful to a buyer, a user, a developer and an auditor.

Use a minimum record with twelve fields: service name, business outcome, accountable owner, technical owner, user group, data classes, connected tools, permitted actions, human approval points, evaluation measures, current lifecycle stage and next review date. Add links to the runbook, model and prompt versions, supplier terms, data protection assessment and incident route rather than copying those documents into the catalogue. Record dependencies too. If a finance agent relies on a particular ERP connector and model endpoint, a change to either can affect the service even when the agent's own code is unchanged.

The leading misconception is that this creates paperwork before value. A good catalogue entry should take less than an hour for a small pilot because most fields are questions the team must answer anyway. If the owner, users, actions and success measures cannot be stated briefly, the experiment is not ready for broader access. The catalogue exposes ambiguity early, when it is still cheap to correct.

Reuse needs an explicit gate or duplication will keep winning

Teams often duplicate agents because building a fresh demonstration feels faster than finding, understanding and adapting an existing service. That behaviour is rational when there is no trusted catalogue. Search costs are high, ownership is unclear and nobody knows whether reuse will create a support obligation. Leaders cannot solve this with a request to collaborate more. They need to make reuse the easier path.

The Department for Work and Pensions offers a useful signal. In September 2026, PublicTechnology reported that DWP had awarded a one-year contract worth £688,630 to support Copilot agents and other digital workplace AI solutions. The stated aim was a scalable, secure and consistent approach to delivery across a department serving 99,000 staff. That is not proof that every organisation needs a central platform or a large consultancy contract. It is evidence that once agent use moves beyond isolated trials, consistency becomes a service-design problem.

Put a reuse checkpoint at the start of intake. The proposing team should identify the business outcome, search the catalogue and compare any close matches across data, actions, users and service levels. There are then four legitimate decisions. Reuse an existing agent unchanged. Extend an existing service for a new user group. Fork it because the risk or data boundary is materially different. Build a new agent because no suitable service exists. Record the decision and the reason.

What this means in practice is that a marketing research agent and a procurement research agent may share retrieval, source-checking and summarisation components while keeping separate data and approval policies. Reuse does not mean forcing every workflow through one giant assistant. It means reusing proven capabilities and controls where the operating context supports it. Track duplicate capability requests, avoided build effort and adoption of shared components. Those measures show whether the catalogue is changing behaviour rather than becoming a static directory.

Make ownership and shutdown conditions visible before launch

An agent can have a product owner, technical owner, data owner and risk owner, yet still have nobody clearly responsible when its output harms a customer or interrupts a process. The catalogue should not reproduce a complicated responsibility matrix. It should identify one accountable service owner and make the supporting roles visible. That owner accepts the operating boundary, reviews evidence and decides whether the agent continues, changes or stops.

The need for practical authority is visible in the same Harness research. While 76% of respondents believed they could disable a misbehaving agent within 15 minutes, only 33% had an instant kill switch. Seventy-four per cent believed their testing would catch a production-impacting failure, but only 19% had a gate that automatically blocked every bad release. These numbers do not mean every SME needs an expensive agent management suite. They show why an inventory without an operating response is insufficient.

Every catalogue entry should state how to pause the service, who can authorise that action and what happens to work already in progress. Define triggers in business language: suspend if the agent exceeds an agreed error threshold, accesses an unapproved data class, performs an action outside its authority, loses a critical dependency or creates an unexplained cost spike. Link to the technical procedure and test it at a frequency proportionate to the risk. A read-only internal summariser might receive a quarterly review. An agent that changes customer records may need a tested shutdown route before every material release.

Do not confuse a stop mechanism with observability. A button is useful only when the organisation can detect that it should be pressed. Record the measures that expose abnormal behaviour, including tool errors, approval rejection rates, exception volumes, latency, spend and business outcome drift. The operating evidence should include trace data rather than relying on chat transcripts alone, as explained in our guide to treating every agent run like a trace. The catalogue gives staff the map; monitoring tells them where to look.

Platform products are useful, but the operating model comes first

The market is responding to agent sprawl with control planes, discovery tools and lifecycle platforms. In September 2026, WSO2 announced general availability of an open-source Agent Manager that separates governance infrastructure from agent logic and adds controls for identity, MCP access, sandboxing, evaluation and observability. The announcement reported by IT News Africa is one example of an emerging product category built to manage agents across different frameworks, models and deployments.

Microsoft's September security update points in the same direction. Microsoft said agents were running across employee devices, cloud platforms and developer workflows, and described controls to discover them, govern reachable resources and contain problems. Its September 2026 security release also made network policies available for human actions and on-behalf-of agent traffic, allowing sensitive transfers to risky destinations to be blocked before data leaves. These developments matter because agent management is moving from a feature inside individual builders towards an estate-level discipline.

Buying a control plane before defining the catalogue will not fix ownership or duplication. Software can discover an endpoint, model call or MCP server. It cannot reliably decide which business outcome the service owns, whether two agents are legitimate variants or which team should fund support. Start with the minimum record and intake gate, then automate the parts that create repeated manual work. For a small estate, a controlled list, named owner and monthly review may be enough. As the estate grows, integrate discovery, identity, evaluation, cost and deployment data.

The counterargument is that a central catalogue will slow builders and force one platform. It should do neither. Make the catalogue framework-neutral and provide a fast lane for read-only, low-risk experiments. Require deeper evidence as actions, sensitive data and user reach increase. The goal is federated delivery with shared visibility, not central approval of every prompt edit.

Launch the catalogue as a 30-day operating change

Do not begin with a six-month taxonomy project. Run a 30-day exercise that produces a usable view of the current estate and changes how the next agent is approved. In week one, name an executive sponsor and a working owner from operations, technology or transformation. Agree the twelve required fields, the lifecycle labels and a simple risk tier based on data sensitivity, action authority and user reach.

In week two, discover what already exists. Interview functional leads, review Copilot and other builder environments, examine API and identity logs where available, and ask procurement for AI-related contracts. Include pilots, personal automations that affect business work, vendor-embedded agents and retired services whose credentials may still exist. Do not advertise the exercise as a hunt for unauthorised tools. Teams disclose more when the purpose is to create a safe route to reuse and support.

In week three, review each entry with its proposed owner. Merge duplicates only where the business boundaries genuinely match. Add missing shutdown conditions, measures and review dates. Select two or three common capabilities, such as approved internal search, document classification or meeting follow-up, that could become reusable components. In week four, introduce the reuse checkpoint for every new request and publish a short route for adding, changing and retiring a service. Give staff searchable summaries but restrict sensitive technical details and credentials.

Measure the catalogue through decisions, not completeness theatre. Useful indicators include the percentage of active services with a named owner and tested stop route, the number of proposals that reuse an existing component, time from request to approved pilot, duplicate capabilities retired and services overdue for review. The Harness survey found that only 53% of agent-related changes passed through any standard pipeline before production and 58% of organisations reported more production incidents per 100 changes after adopting agents. A catalogue will not solve release quality by itself, but it creates the shared object that an evaluation, approval and deployment process can govern.

After 30 days, leaders should be able to answer what exists, what can be reused, where the riskiest gaps sit and who can act. That is enough to stop agent sprawl becoming the default operating model while preserving room for teams to experiment.

Frequently Asked Questions

What is an AI agent service catalogue?

It is a searchable record of the AI agents used by an organisation, described as business services. It captures purpose, ownership, users, data, tools, permitted actions, controls, evidence and lifecycle status.

How is a service catalogue different from an AI inventory?

An inventory proves that a technical asset exists. A service catalogue also explains the outcome it supports, who is accountable, how it is operated and whether another approved service can be reused.

Does a small UK business need agent management software?

Usually not at first. A controlled list, clear owners, a reuse checkpoint and tested shutdown routes can cover a small estate. Add automated discovery and lifecycle tooling when manual reviews become unreliable.

Who should own the AI agent catalogue?

Choose one operational owner with authority to maintain the process, supported by technology, security, data protection and functional leaders. Each listed agent should still have one accountable service owner.

Will a central catalogue slow down experimentation?

It should speed up low-risk work by making approved components easier to find. Use a light entry for read-only experiments and require deeper evidence only as data sensitivity, autonomy and user reach increase.

What should trigger retirement of an AI agent?

Retire an agent when its business need ends, adoption stays below an agreed threshold, a better shared service replaces it, controls cannot be maintained or its cost and failure rate no longer justify operation.

How often should catalogue entries be reviewed?

Set review frequency by risk. Quarterly may suit low-risk internal tools, while agents that act on customer, employee or financial records should be reviewed after material changes and on a shorter fixed cycle.