Skip to content
Search
paperAugust 2026Unreviewed

How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation

Kimberly Milner, Minghao Shao, Nanda Rani, Haoran Xi, Venkata Sai Charan Putrevu, Meet Udeshi, Sandeep K. Shukla, Prashanth Krishnamurthy, Farshad Khorrami, Muhammad Shafique, Ramesh Karri

Abstract

Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with direct flag exposure, memorized recall, external lookup, guessing, and unsupported claims, potentially overstating the agent's cybersecurity capability. We introduce CTF-ABACUS, a trace-based agent audi

Categories

Cite

@misc{milner2026how,
  title = {{How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation}},
  author = {Kimberly Milner and Minghao Shao and Nanda Rani and Haoran Xi and Venkata Sai Charan Putrevu and Meet Udeshi and Sandeep K. Shukla and Prashanth Krishnamurthy and Farshad Khorrami and Muhammad Shafique and Ramesh Karri},
  year = {2026},
  month = aug,
  eprint = {2608.26237},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.26237}
}