May 2026Unreviewed
LLM-Agnostic Semantic Representation Attack
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Tairan Huang, Shaohui Mei, Lap-Pui Chau
Abstract
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting adversarial prompts. Predominant token-level optimization methods primarily rely on optimizing for exact affirmative templates (e.g., ``\textit{Sure, here is...}''). However, these paradigms frequently encounter bottlenecks such as suboptimal convergence, compromised prompt naturalness, and poor cross-model generalization. To addres
Categories
Cite
@misc{lian2026llmagnostic,
title = {{LLM-Agnostic Semantic Representation Attack}},
author = {Jiawei Lian and Jianhong Pan and Lefan Wang and Yi Wang and Tairan Huang and Shaohui Mei and Lap-Pui Chau},
year = {2026},
month = may,
eprint = {2605.08898},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.08898}
}