Skip to content
Search
paperAugust 2026Unreviewed

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen

Abstract

Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails. We present AGENTCHAOSBENCH, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry. We run five heterogeneous applications that coordinate agents over the Agent-to-Agent protocol and call tool

Categories

Cite

@misc{zhang2026whenb,
  title = {{When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry}},
  author = {Chenkai Zhang and Yiran Li and Yifang Tian and Michalis Bachras and Hans-Arno Jacobsen},
  year = {2026},
  month = aug,
  eprint = {2608.14680},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/12ee9654cd5677de7f8cd2ecff7ad773b3f779a8}
}