← Back to News
ANALYSIS

The Agent-Observability Turn — Gemini Spark Meets OpenTelemetry's MCP Spans

In one week Google shipped a 24/7 autonomous agent and OpenTelemetry shipped tool-layer spans for agent traces. The same week the EU AI Act's logging clock ticks toward August. Autonomous agents just became something you have to be able to watch, and the standards to watch them are arriving on cue.

By Michael Eakins•• min read
AI AgentsObservabilityOpenTelemetryModel Context ProtocolGoogle Gemini

Executive Summary

Two announcements landed within days of each other this week, and on the surface they have nothing to do with each other. Google made Gemini 3.5 Flash generally available and introduced Gemini Spark, a personal agent meant to run around the clock and take actions on a user's behalf. Separately, the OpenTelemetry project shipped Model Context Protocol spans — tracing instrumentation that finally reaches into the tool layer an agent calls, rather than stopping at the model request. Put them side by side and the connection is obvious. The industry spent two years making agents that act autonomously. It is now scrambling to make those agents something an operator can actually watch, reconstruct, and audit after the fact. That second project does not have a launch event, but it is the one that decides whether autonomous agents are deployable in any setting that carries real consequences.

The News

Three threads converged this week, and reading them together is the whole point.

Google shipped an always-on agent. Gemini 3.5 Flash went generally available through Google Antigravity and the Gemini API, positioned as frontier-class intelligence at Flash-tier latency and cost, and reported to beat the previous Pro model on coding and agentic benchmarks. Alongside it, Google introduced Gemini Spark — described as a 24/7 personal agent that operates autonomously under user direction and checks in before major actions. The design pattern matters more than the product name. An agent that runs continuously, holds goals across sessions, and calls tools on your behalf is no longer a single request you can read in a log. It is a long-running process whose behavior is distributed across dozens of model calls and tool invocations, most of which you will never see unless you instrumented for them in advance.

OpenTelemetry reached the tool layer. In the same window, OpenTelemetry's generative-AI instrumentation added Model Context Protocol spans. The significant detail is that MCP spans enrich the existing tool-execution span rather than duplicating it — so an agent trace can now show the model decision and the tool call it triggered as parent and child in one continuous tree, instead of a model span that simply goes dark the moment the agent reaches for a tool. That tool layer was the black box. An agent would decide to call a function, and the trace ended there; whether the call succeeded, what it returned, how long it took, and whether it was even the right tool were invisible. Closing that gap is what makes an agent trace a debugging artifact instead of a partial transcript.

The regulatory clock kept ticking. The EU AI Act's Article 12 requires high-risk AI systems to keep automatic logs sufficient to reconstruct individual AI-assisted decisions after the fact, with broader enforcement obligations landing in August 2026. "Reconstruct the decision" is, functionally, a specification for a trace. A system that cannot show what it did, in what order, on whose input, with which tool, is a system that cannot satisfy that requirement — no matter how good its outputs look on a demo.

What changed this week

The tool layer

Agent traces could already show the model request. The new MCP spans extend the same trace into the tool calls the model triggers — turning a partial transcript into a reconstructable record of what the agent actually did.

↑ 1%black box closed: model-decision-to-tool-call is now one continuous span tree

Deep Dive

Technical Implications

The reason agent observability is hard is that an agent is not one call. A single user request to something like Gemini Spark can fan out into a planning step, several tool calls, a few follow-up model calls to interpret the tool results, and a final synthesis — each of which can fail in its own way. The model can hallucinate a tool that does not exist. The tool can return an error the model silently ignores. The agent can loop, re-calling the same tool with the same bad arguments until it burns its step budget. None of these are visible in a metric that only counts requests and tokens.

A trace is the right shape for this problem because it is a tree, and an agent run is a tree. The root span is the invocation. Child spans are the model calls and the tool calls, nested in the order they happened, each carrying its own attributes — which model, how many tokens, which tool, what arguments, success or error. When something goes wrong, you do not read a wall of logs; you open the trace and see exactly which branch turned red. That is why the OpenTelemetry generative-AI conventions, which standardize the gen_ai.* attributes on model spans and now the MCP attributes on tool spans, are quietly more important than any single vendor dashboard. They define the schema everyone else builds on.

Where agent behavior was already visible vs. where it went dark (illustrative)

Where agent behavior was already visible vs. where it went dark (illustrative)
layervisibility
Tool calls (was untraced)18
Model decisions71
Token + cost metrics88
Final output only97

The practical upshot for engineers is that instrumenting an agent is now a standards exercise, not a bespoke one. You wrap the agent invocation in a root span, wrap each model call and each tool call in child spans, set the conventional attributes, and export to any backend that speaks OpenTelemetry. Because the conventions are shared, the same instrumented agent emits traces that Datadog, Grafana, an open-source collector, or a self-hosted Jaeger can all read without custom adapters. That portability is the difference between observability you own and observability you rent. For a step-by-step build, see this week's companion tutorial on instrumenting an MCP tool-use agent with OpenTelemetry in TypeScript, which constructs exactly this span tree from a root agent.invoke down through the model and tool spans.

Business Impact

A market has already formed around this gap. The agent-observability tooling landscape for 2026 includes Braintrust, LangSmith, Arize Phoenix, Helicone, Galileo, Maxim, and Datadog's LLM observability product, among others — a clear signal that "watch the agent" is now a budget line, not a research curiosity. The split worth watching is between platforms that lock you into a proprietary trace format and those that commit to OpenTelemetry portability, because that choice determines whether switching vendors later means re-instrumenting every agent you run.

The same agent, with and without instrumentation

Blind autonomous agent

Failure modeSilent — surfaces as a bad outcome days later
DebuggingRe-run and hope it reproduces
AuditNo record of which tool acted on what
Vendor lock-inBehavior trapped in one vendor console

Traced autonomous agent

Failure modeVisible — the red span names the failing tool call
DebuggingOpen the trace, read the tree
AuditFull reconstructable record per decision
Vendor lock-inPortable OpenTelemetry traces, any backend

There is also a cost story. Agents are expensive precisely because they loop, and an untraced loop is a bill you discover after it has already run. Traces make the fan-out legible: you can see that a task took eleven tool calls when it should have taken three, and fix the prompt or the tool description that caused it. The teams treating observability as optional are the ones who will be surprised by both their incident reports and their invoices.

Industry Implications

The deeper shift is cultural. For most of the last two years, the agent conversation was about capability — can it plan, can it use tools, can it run unattended. Gemini Spark is an answer to that question: yes, it can run unattended. The very fact that the answer is now yes is what moves the hard problem downstream. An agent you supervise turn-by-turn does not strictly need a trace, because you are the trace. An agent that runs at 3 a.m. while you sleep needs to be able to tell you, in the morning, exactly what it did — and "exactly what it did" is a span tree.

This is also why the standards timing is not a coincidence. Autonomy and observability are two halves of the same product requirement, and the market tends to ship the exciting half first and the necessary half second. We saw the same sequence with microservices a decade ago: distributed systems went mainstream, everything became impossible to debug, and distributed tracing — OpenTelemetry's own origin story — emerged to make them operable again. Agents are the new distributed system, except the nondeterminism is inside the nodes. The tooling is recapitulating that arc on fast-forward.

From capable agents to observable agents

2024

Agents go mainstream

Tool-using, multi-step agents move from demos into products; observability is an afterthought.

2025

The blind-agent problem

Teams discover that token metrics and output logs cannot explain why an autonomous run went wrong.

Jun 2026

MCP spans + always-on agents

OpenTelemetry reaches the tool layer the same week Google ships a 24/7 autonomous agent.

Aug 2026

Regulatory logging pressure

EU AI Act obligations push reconstructable per-decision logging from nice-to-have toward mandatory for high-risk systems.

2027 (forecast)

Conventions stabilize

GenAI and MCP semantic conventions reach Stable and ship default-on in major agent frameworks.

If the conventions stabilize and ship by default inside the major frameworks, instrumentation stops being a thing engineers remember to add and becomes a thing that is simply on — the way HTTP request tracing is on in a modern web framework without anyone thinking about it. That is the outcome I am betting on in my prediction that OpenTelemetry's GenAI and MCP conventions reach Stable and ship default-on in major agent frameworks by the end of 2027. The same operability discipline applies when you move from one agent to many; the coordination and failure surface only grows, as covered in the build guide for a parallel subagent orchestrator in TypeScript.

Conclusion

The week's two headlines are really one story told from both ends. Gemini Spark is the demand: agents are now autonomous enough that nobody can sit and watch them work. OpenTelemetry's MCP spans are the supply: a standard, portable way to let the agent show its work after the fact. The regulatory clock is the deadline that turns a best practice into a requirement. The teams that internalize this now — instrumenting the agent loop down to the tool call, exporting portable traces, treating the span tree as the primary debugging surface — are the ones who will be able to run autonomous agents where it matters. The teams that ship the autonomy and skip the observability are building systems they will not be able to explain, debug, or defend the first time one of them does something surprising at 3 a.m.

Sources