September 2026Unreviewed
IC-GCG: Jailbreaking Large Language Models via Intermediate Consistency Optimization
Zichu Ren, Donghai Zhu, Haibo Hong, Jun Shao
IEEE Internet of Things Journal
Abstract
Recent jailbreak attacks demonstrate that large language models (LLMs) can be manipulated to generate harmful outputs through adversarial prompts even after robust alignment. However, prevailing methods typically focus on forcing a desired response at the output layer—a surface-level strategy that is brittle and often fails to bypass the more fundamental safety checks embedded within the model’s internal mechanisms. In contrast, we propose intermediate consistency greedy coordinate gradient (IC-
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{ren2026icgcg,
title = {{IC-GCG: Jailbreaking Large Language Models via Intermediate Consistency Optimization}},
author = {Zichu Ren and Donghai Zhu and Haibo Hong and Jun Shao},
year = {2026},
month = sep,
journal = {IEEE Internet of Things Journal},
doi = {10.1109/JIOT.2026.3707160},
url = {https://www.semanticscholar.org/paper/7ecc8bf212ba04a3db3043938853af5bca5e6533}
}