Skip to content
Search
paperJune 2026Unreviewed

NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense

Zhongyang Lin, Ziran Zhao, Feifei Zhai, Pengyuan Liu

Abstract

Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translation, encoding, adversarial suffixes, and multi-turn buildup. Existing defenses still struggle to handle these attacks without over-blocking benign but sensitive requests, partly because they often apply the same action to every prompt and therefore fail to balance safety and helpfulness. We propose NeuroArmor, a white-box runtime defense that uses pr

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{lin2026neuroarmor,
  title = {{NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense}},
  author = {Zhongyang Lin and Ziran Zhao and Feifei Zhai and Pengyuan Liu},
  year = {2026},
  month = jun,
  eprint = {2606.03486},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.03486}
}