Agent Observability Is Now a Buying Requirement: What to Demand Before Agents Touch Production Data
An agentic feature that cannot be traced is not a feature. It is an account sitting inside your data-path that nobody owns. Over the past two quarters, the technical-buyer conversation has shifted: the question is no longer whether a vendor’s AI agent is impressive, but whether you can see what it did, explain why, and undo it. Vendors without an answer are starting to lose deals on this criterion alone.
This piece sets out the four observability requirements we now treat as table stakes when a vendor proposes to run agents against production data, and how to test each claim before signing.
The four requirements
| Requirement | What it must contain | The test that exposes an empty claim |
|---|---|---|
| Tracing | Per-run trace: prompt, tools called, inputs, outputs, latency, token cost | Ask the vendor to show you a trace from a rival customer’s sandbox, not a demo slide |
| Evals | A named evaluation suite the vendor runs on every release, with failure thresholds | Ask for the eval run history. If it does not exist, the eval does not exist |
| Guardrails | Deterministic checks (allowlists, scope limits, rate caps) that run outside the model | Ask which guardrail fires when the agent misbehaves, and who wrote it |
| Audit logging | Immutable record of every action taken against your tenant, exportable to your SIEM | Ask for a 30-day export sample. Check field coverage, not format |
A vendor that passes one of these and fails the rest is not offering observability. It is offering a screenshot of observability.
Tracing: run-level, not request-level
Traditional APM gives you request traces. A single HTTP call in, a response out. Agent systems break that model. One user action can fan out into dozens of model calls, tool invocations and sub-agent handoffs, several minutes apart.
The requirement is a run-level trace. One identifier that follows the entire agent execution from trigger to completion, with each model call and tool call hanging off it. The trace must carry the exact prompt sent to the model, not a template with the variables stripped. It must record which model version answered, because the same vendor agent can route across versions mid-day.
Ask specifically how long traces are retained and at what plan tier. Retention below 90 days is useless for any incident you discover late. Several vendors sell tracing as a premium-tier addon. That pricing choice tells you whether the vendor treats observability as a product feature or an enterprise tax.
Evals: the release-gate question
Every serious agent vendor claims to evaluate. Almost none will show you the evaluation when asked directly.
The demand to make is simple: a per-release eval report on a named suite, with a threshold that fails the build. If a vendor cannot name the suite, state the pass threshold, or describe what regressed in the last release and what changed as a result, the eval programme is marketing.
The evals that matter are task-completion evals on representative production-shaped inputs, plus regression evals pinned against the previous release. Vibe checks and offline benchmark scores do not count; they rarely predict behaviour under your data distribution. Expect to push hard here. Vendors are not accustomed to showing failure numbers, and the discussion will tell you more about engineering culture than the numbers themselves.
Guardrails: code, not prompts
A guardrail written into the system prompt is a suggestion. The model can ignore it and increasingly does under adversarial input or long-context drift.
The structural test: a guardrail is credible only when enforced in code outside the model. Concretely, that means schema-validated tool-call arguments, hard allowlists on which tables or APIs the agent can write to, and execution-time rate caps that trip regardless of what the model decided. Ask the vendor to point at the line of code. If the answer is a prompt excerpt, the guardrail is not a guardrail.
Also ask what happens when a guardrail trips. Does the run halt, degrade to human review, or silently continue with reduced scope? The right answer depends on the use case, but the vendor should be able to describe the behaviour precisely. Vague answers here are a red flag for the entire safety architecture.
Audit logging: built for your SOC
If an agent acts on your behalf, the log of its actions belongs in your SIEM, not the vendor’s dashboard. The demand is three properties.
- Immutability. The vendor cannot rewrite or selectively delete entries after the fact. Hash-chained or write-once storage is the credible pattern.
- Field coverage. Every action carries tenant, actor (human or agent), run identifier, resource touched, before/after state for writes, and timestamp.
- Exportability. Structured, streaming export to your SIEM of choice, not a quarterly CSV on request.
Ask for a sample export covering a failure case, not a happy path. The failure case is where vendors reveal that certain action types never made it into the log at all.
The bottom line for buyers
Agent observability is no longer an aspirational roadmap item. It is a contract term. Vendors who cannot meet the four requirements are not saying “we will get there.” They are saying they have not yet built the systems that would let you trust them on your production data, and you should price that accordingly.
Put the demands in the evaluation questionnaire now. The vendors who survive that scrutiny are the ones whose agentic features you can actually run.