Skip to content
Search
paperAugust 2026Unreviewed

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, Sungheon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani

Abstract

Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments

Categories

Cite

@misc{masukawa2026trident,
  title = {{Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)}},
  author = {Ryozo Masukawa and Ian Bryant and Armita Kazeminajafabadi and Sanggeon Yun and Hyunwoo Oh and Sungheon Jeong and Nathaniel D. Bastian and Mahdi Imani and Mohsen Imani},
  year = {2026},
  month = aug,
  eprint = {2608.04317},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/f113cf61f49b1f85a1e9dc20d3d4e42a00917f68}
}