August 2026Unreviewed
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
Hujian Zhu, Yihao Huang, Felix Juefei-Xu, Xinfeng Li, Peng Zeng, Simeng Qin, Qing Guo, G. Pu
Abstract
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass explicit safety mechanisms by replacing harmful terms in original harmful questions with benign alternatives and leveraging contextual information to induce the target model to reinterpret these alternatives as their corresponding harmful concepts. However, existing sem
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhu2026ico,
title = {{ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization}},
author = {Hujian Zhu and Yihao Huang and Felix Juefei-Xu and Xinfeng Li and Peng Zeng and Simeng Qin and Qing Guo and G. Pu},
year = {2026},
month = aug,
eprint = {2608.03210},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/f6f8fcc18350b6b6d2f859b383e0327b136d7d93}
}