April 2026Unreviewed
Double-Gaming: Jailbreak Attacks against LLM based on Red-Blue Team Game Theory
Chenlu Ma, Huairui Zhao, G. Nie, Baiyang Ji, Beibei Li, Haibin Zheng
2026 8th International Conference on Software Engineering and Computer Science (CSECS)
Abstract
Despite existing security alignment mechanisms, Large Language Models (LLMs) remain vulnerable to jailbreak attacks under static defenses. To address this, we propose a novel jailbreak attack and defense optimization framework based on Red-Blue Team dynamic game theory. This framework establishes an automated adversarial mechanism between red and blue teams, achieving closed-loop optimization of instruction generation, semantic interception, and strategy evolution. The red team employs reinforce
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@inproceedings{ma2026doublegaming,
title = {{Double-Gaming: Jailbreak Attacks against LLM based on Red-Blue Team Game Theory}},
author = {Chenlu Ma and Huairui Zhao and G. Nie and Baiyang Ji and Beibei Li and Haibin Zheng},
year = {2026},
month = apr,
booktitle = {2026 8th International Conference on Software Engineering and Computer Science (CSECS)},
doi = {10.1109/CSECS69124.2026.11541892},
url = {https://www.semanticscholar.org/paper/17d55e7e6983830e40b9cb742d612e16687b8365}
}