August 2026Unreviewed
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi, Sanggeon Yun, Hyunwoo Oh, Sungheon Jeong, Nathaniel D. Bastian, Mahdi Imani, Mohsen Imani
Abstract
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments
Categories
Cite
@misc{masukawa2026trident,
title = {{Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)}},
author = {Ryozo Masukawa and Ian Bryant and Armita Kazeminajafabadi and Sanggeon Yun and Hyunwoo Oh and Sungheon Jeong and Nathaniel D. Bastian and Mahdi Imani and Mohsen Imani},
year = {2026},
month = aug,
eprint = {2608.04317},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/f113cf61f49b1f85a1e9dc20d3d4e42a00917f68}
}