August 2026Unreviewed
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi, Sandeep K. Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Muhammad Shafique, Ramesh Karri
Abstract
Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with direct flag exposure, memorized recall, external lookup, guessing, and unsupported claims, potentially overstating the agent's cybersecurity capability. We introduce CTF-ABACUS, a trace-based agent audi
Categories
Cite
@misc{milner2026how,
title = {{How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation}},
author = {Kimberly Milner and Minghao Shao and Nanda Rani and Haoran Xi and Venkata Sai Charan Putrevu and Meet Udeshi and Sandeep K. Shukla and Prashanth Krishnamurthy and Farshad Khorrami and Muhammad Shafique and Ramesh Karri},
year = {2026},
month = aug,
eprint = {2608.26237},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.26237}
}