89% of AI agent teams use observability, but fewer than half run evals

A LangChain survey of 1,340 AI practitioners, published in May 2026, found that 89% had implemented observability for their agents, yet only 52.4% ran offline evaluations and 37.3% ran online evaluations. Fewer than one in three teams ran both types of evaluation. The core problem is that observability tools — which track latency, token usage, and span status — cannot detect whether an agent's output was actually correct. A confidently wrong action, such as refunding the wrong order, produces identical telemetry to a correct one, with all spans showing green and no errors logged. Despite this gap, the same survey identified output quality as the leading barrier to deploying agents in production.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in