August 2026Unreviewed
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen
Abstract
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails. We present AGENTCHAOSBENCH, a benchmark for detecting and localizing runtime faults in agentic systems from their execution telemetry. We run five heterogeneous applications that coordinate agents over the Agent-to-Agent protocol and call tool
Categories
Cite
@misc{zhang2026whenb,
title = {{When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry}},
author = {Chenkai Zhang and Yiran Li and Yifang Tian and Michalis Bachras and Hans-Arno Jacobsen},
year = {2026},
month = aug,
eprint = {2608.14680},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/12ee9654cd5677de7f8cd2ecff7ad773b3f779a8}
}