August 2026Unreviewed
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
Xuyang Liu, Yibin Han, Zhenwei Zhang, Kai Chang, Zhiwei Xu, Tian Qiu, Weixian Deng, Jiabao Gao, Xiaolin Peng, Hai Wan, Xibin Zhao
Abstract
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions. However, existing benchmarks mainly evaluate final outputs or aggregate accuracy, providing limited insight into how errors arise and propagate across intermediate reasoning stages. We present DiagChain, a diagnostic benchmark for evidence-grounded attack chain reconstruction that enables stage-wise evaluation of LLM
Categories
Cite
@misc{liu2026diagchain,
title = {{DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction}},
author = {Xuyang Liu and Yibin Han and Zhenwei Zhang and Kai Chang and Zhiwei Xu and Tian Qiu and Weixian Deng and Jiabao Gao and Xiaolin Peng and Hai Wan and Xibin Zhao},
year = {2026},
month = aug,
eprint = {2608.03591},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.03591}
}