June 2026Unreviewed
NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense
Zhongyang Lin, Ziran Zhao, Feifei Zhai, Pengyuan Liu
Abstract
Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translation, encoding, adversarial suffixes, and multi-turn buildup. Existing defenses still struggle to handle these attacks without over-blocking benign but sensitive requests, partly because they often apply the same action to every prompt and therefore fail to balance safety and helpfulness. We propose NeuroArmor, a white-box runtime defense that uses pr
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{lin2026neuroarmor,
title = {{NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense}},
author = {Zhongyang Lin and Ziran Zhao and Feifei Zhai and Pengyuan Liu},
year = {2026},
month = jun,
eprint = {2606.03486},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.03486}
}