May 2026Unreviewed
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang, Zhiyi Tian, Feng Wu, Ahmed Asiri, Shui Yu
Abstract
Representation engineering (RepE) defenses have shown strong robustness against jailbreak attacks on large language models (LLMs). However, these methods fundamentally rely on black-list supervision: they learn jailbreak-to-refusal activation transformations from harmful or jailbreak data that are inherently incomplete and continuously evolving. Hence, the performance of RepE-based defenses becomes tightly coupled to the quality and coverage of collected harmful samples, leaving models vulnerabl
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{chen2026ellipsoid,
title = {{Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling}},
author = {Luoyu Chen and Weiqi Wang and Zhiyi Tian and Feng Wu and Ahmed Asiri and Shui Yu},
year = {2026},
month = may,
eprint = {2605.24552},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.24552}
}