Treat Every AI Agent Run Like a Trace, Not a Transcript
Tools & Technical Tutorials
6 October 2026 | By Ashley Marshall
Quick Answer: Treat Every AI Agent Run Like a Trace, Not a Transcript
UK businesses should instrument AI agent runs as end-to-end traces, with separate spans for model calls, retrieval, tool use and approvals. Start with metadata, route it through OpenTelemetry and keep sensitive prompt content off by default.
A transcript tells you what an AI agent said. A trace tells you what it touched, what failed, what it cost and whether somebody should stop it.
A transcript is not an operational record
Most teams begin monitoring an AI assistant by saving its prompt and final answer. That is useful for reviewing tone or factual quality, but it is a poor record of an agentic workflow. An agent can search a knowledge base, call a customer system, retry an API, ask another model for a judgement and write to a live service before it produces one polite paragraph. The transcript may show none of those steps clearly. If the result is wrong, slow or unexpectedly expensive, the operator is left reconstructing the journey from several unrelated logs.
The better mental model comes from ordinary distributed systems. Treat one agent run as a trace. Give the run a trace identifier, then record retrievals, model calls, tool executions, policy checks and human approvals as child spans. Each span should carry the small set of facts needed to answer operational questions: which model or tool ran, how long it took, whether it failed, how many tokens it consumed and which policy decision allowed it. This turns a vague report that the agent behaved oddly into a testable sequence of events.
This is also where current UK guidance is heading. The National Cyber Security Centre's August 2026 advice says operators need reliable access to telemetry both in near real time and afterwards. It recommends logging agent activity alongside events from the sandbox, proxies and network environment, then protecting those records from modification or deletion. That is stronger than keeping a chat history. It makes observability part of the security control, not merely a developer convenience.
What this means in practice is simple: do not approve an autonomous workflow because the demonstration looked sensible. Ask whether a single customer outcome can be followed from request to response across every model, tool and approval boundary. If the answer is no, the organisation cannot yet explain the system it is operating.
Use a shared vocabulary before buying another dashboard
AI observability is becoming crowded with specialist products, but the first decision should be about the data format rather than the dashboard. OpenTelemetry semantic conventions provide common names for traces, metrics, logs and resources. The emerging generative AI conventions extend that approach to model and agent operations. Instead of every framework inventing incompatible fields, teams can describe a model request, an agent invocation, a workflow, a retrieval and a tool execution in a consistent way.
The practical value is portability. A span for a tool call can include the tool name, call identifier, duration and error type. A model span can record the requested model, token usage and latency. A retrieval span can identify the data source and result count. Those records can flow through the OpenTelemetry Collector into an existing observability platform rather than forcing the business to create a second monitoring estate just for AI. That matters when an agent error actually starts in a normal dependency such as a CRM API, database or identity service.
A September 2026 technical review from Dash0 notes that the generative AI conventions moved into a dedicated repository in OpenTelemetry version 1.42.0 on 12 June 2026 so they could evolve separately. It also describes operation names for agent creation, agent invocation, workflow invocation, planning, tool execution, retrieval and memory operations. That breadth is useful, but the conventions are still developing. Teams should pin instrumentation versions and record the schema version behind each dashboard.
The counterargument is that a stable internal schema would be safer than adopting an evolving standard. That can be true for a very narrow system. It becomes expensive once the organisation uses several models, frameworks or observability vendors. A better compromise is to use OpenTelemetry at the boundary, keep unstable attribute names behind one adapter and transform older dialects in the Collector. The business gets a common operational language without hard-coding every experimental field throughout its applications.
Trace the decisions that create cost and risk
A useful trace is not a data dump. It is a structured account of the decisions that can change cost, customer impact or exposure. Start with a root span for the business request. Under it, create spans for retrieval, each model call, each tool call, each policy decision and each human checkpoint. Add the workflow version, agent identity and environment. For tools that can change records, transfer money, send messages or expose personal data, record whether execution was attempted, allowed, denied or approved.
The most revealing metrics are often ratios rather than totals. Track model calls and tool calls per successful task. If model calls rise while tool calls stay flat, the agent may be deliberating or looping without acting. If tool calls rise, it may be retrying a failing integration. Token usage by itself cannot tell those stories. Pair it with the completion status, elapsed time and business outcome, then calculate cost per successful outcome rather than cost per million tokens. This connects technical telemetry to the unit economics that a finance team can use.
Operational alerts should reflect the workflow's intended shape. A research assistant might be allowed ten retrievals but no write action. A customer service agent might read several systems, draft a response and require approval before sending. Alert on a write attempt without approval, a tool outside the allowlist, a sudden increase in calls per run, repeated identical parameters or an execution time beyond the normal range. These are concrete signals that can reach the same security and service teams already watching the rest of the platform.
This complements the NCSC's warning that greater autonomy increases potential impact and therefore requires stronger controls. It also builds on the access and identity issues discussed in our guide to AI agent identity lifecycle controls. Identity tells you who or what acted. A trace connects that identity to the exact sequence of actions. Together they turn accountability from a policy statement into evidence.
Keep sensitive content out until you have a reason to capture it
The obvious temptation is to record every prompt, response and retrieved document. That makes debugging easier, but it can create a second, less controlled copy of personal data, confidential material and credentials. Observability systems are usually searchable by many technical users, retained for long periods and exported to third parties. A monitoring improvement can therefore enlarge the privacy and security problem it was meant to solve.
Begin with metadata. Model name, operation type, token counts, duration, error type, tool name, approval result and trace identifier are enough to diagnose a surprising number of problems. The OpenTelemetry generative AI approach deliberately separates message content from ordinary span attributes and treats content capture as opt-in. That is a sound default for UK organisations. Capture full content in a controlled test environment where synthetic or minimised data is available. In production, enable it only for a defined purpose, for a limited period and behind access controls.
Where content is genuinely required, decide what will be redacted before telemetry leaves the application. Remove secrets, authentication headers, personal identifiers and unnecessary document text at source. Define retention separately for metadata and content. Restrict search and export rights, log access to the observability platform and make deletion procedures workable. The NCSC also points out that logs may contain sensitive information and should be protected from modification or deletion. Immutability and privacy are not opposites: records can be tamper-evident while still using short, policy-based retention for sensitive fields.
There is an important misconception here. Some leaders assume that full prompts are necessary to prove what happened. They are not sufficient. A prompt does not prove which credentials were used, what network request followed, whether a policy engine intervened or which database row changed. Metadata from the surrounding system is often stronger evidence. Treat content as a diagnostic extension of the trace, not as the trace itself. This approach supports data minimisation while preserving the operational facts needed for incident response.
Build a minimum viable trace in two weeks
A business does not need a perfect enterprise observability programme before it can improve one workflow. Choose a single agentic process with a clear owner and measurable outcome. Map the root request, the models used, the tools available, any retrieval source and the point where a human approves consequential action. Give each run one trace identifier and pass it through every component. If a vendor API cannot accept that identifier, store its request identifier on the corresponding span so the records can still be joined.
During the first week, instrument metadata only. Record the agent and workflow version, model, tool, timestamps, duration, result, error type, token counts and approval state. Send the data through an OpenTelemetry Collector to the organisation's existing backend. Build three views: a trace waterfall for one run, a daily view of successful and failed outcomes, and a table of the slowest or most expensive runs. Test that a support engineer can move from a customer reference to the relevant trace without searching several systems manually.
During the second week, add alerts and exercises. Force a tool timeout, deny an approval, return empty retrieval results and trigger an unauthorised tool request. Confirm that the trace distinguishes each failure and that the right owner receives the alert. Set an initial ceiling for model calls, tool calls, elapsed time and spend per run. These ceilings do not need to be permanent. Their purpose is to expose unexpected behaviour while the team learns the normal operating range.
Finally, run an incident replay with security, operations and the process owner. Ask them to explain what the agent attempted, which data source it used, which action succeeded and how the run ended. If they cannot answer from the trace, add the missing event. If they can answer only by exposing complete prompt content to everyone, redesign access and redaction. This is the practical test: the telemetry should make the workflow explainable without making sensitive content casually available.
Observability is evidence, not a substitute for control
Tracing an agent does not make the agent safe. It tells you what happened and can help you intervene quickly, but it cannot compensate for excessive permissions, weak sandboxing or an undefined risk appetite. A beautifully instrumented agent with access to every customer record is still badly designed. Observability must sit beside least privilege, allowlisted tools, short-lived credentials, approval gates and a tested shutdown route.
The OWASP GenAI LLM Top 10 2026, published on 3 August, says its updated guidance draws on thousands of real-world AI security incidents and maps risks to frameworks including NIST, MITRE ATLAS and the OWASP Top 10 for Agentic Applications. The business lesson is that monitoring should connect to recognised security practice. A tool invocation that violates policy should not merely appear in a dashboard. It should be denied, recorded and routed into an established response process.
This is also why buying a specialist platform should come after the trace design. Vendors can provide evaluations, prompt analytics and polished visualisations, but they cannot decide which customer outcome matters, which action requires approval or who owns a failure. Those are operating model decisions. Define them first, then test whether existing OpenTelemetry infrastructure can carry the required evidence. Add specialist tooling only where it closes a demonstrated gap, such as automated quality evaluation or sensitive-data redaction.
For UK leaders, the board-level question is not whether every technical field has been captured. It is whether the organisation can detect a material deviation, stop harmful activity, investigate the sequence and show what changed afterwards. A modest trace with reliable identities, policy outcomes and protected logs is more valuable than an enormous transcript archive nobody can interpret. Before the next agent moves into production, require one successful incident replay. If the team can explain and contain the failure, observability is doing useful work. If it can only display the final answer, the system is still operating in the dark.
Frequently Asked Questions
What is an AI agent trace?
It is an end-to-end record of one agent run. The root trace links child spans for model calls, retrievals, tool actions, approvals and policy decisions so operators can reconstruct what happened.
Is a saved prompt and response enough for auditing?
No. It rarely proves which tool ran, which credentials were used, what system changed or whether an approval gate was applied. System and policy telemetry must be linked to the conversation.
Do we need a specialist AI observability platform?
Not necessarily. Start by exporting standard OpenTelemetry data into the observability platform you already operate. Add specialist tooling only for a clear gap such as automated evaluation, redaction or prompt analysis.
Should production traces include full prompts and responses?
Usually not by default. Start with operation metadata, timings, token counts, tools, errors and policy results. Enable content capture only for a defined purpose with redaction, restricted access and limited retention.
Which metrics reveal an agent loop?
Track model calls and tool calls per run, elapsed time, repeated parameters and cost per outcome. Rising model calls with flat tool activity often indicates deliberation without action, while repeated tool calls suggest retries or a stuck integration.
How should traces handle personal data under UK GDPR?
Apply data minimisation, purpose limitation, access controls and defined retention. Redact at source where possible, separate content from metadata and document why any production content capture is necessary.
How long does a first implementation take?
A focused team can instrument and exercise one bounded workflow in about two weeks. Enterprise-wide standards, retention and integration will take longer, but they should follow evidence from that pilot.
Does observability replace sandboxing and approval gates?
No. It provides detection and evidence. Least privilege, sandboxing, allowlisted access, approval controls, short-lived credentials and tested shutdown procedures still prevent or limit harm.