Skip to content
Search
paperSeptember 2026Unreviewed

Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning

Tejasvi C. Addagada

Abstract

Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task completed by pattern rather than judged as content; and the erosion of safety alignment outside English. A natural hypothesis is that these compound. We test it directly. Using a deterministic IICL operator and a StrongREJECT-style rubric judge, we red-team two

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{addagada2026structural,
  title = {{Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning}},
  author = {Tejasvi C. Addagada},
  year = {2026},
  month = sep,
  eprint = {2609.08373},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.08373}
}