Skip to content
Search
paperMay 2026Unreviewed

LLM-Agnostic Semantic Representation Attack

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Tairan Huang, Shaohui Mei, Lap-Pui Chau

Abstract

Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting adversarial prompts. Predominant token-level optimization methods primarily rely on optimizing for exact affirmative templates (e.g., ``\textit{Sure, here is...}''). However, these paradigms frequently encounter bottlenecks such as suboptimal convergence, compromised prompt naturalness, and poor cross-model generalization. To addres

Categories

Cite

@misc{lian2026llmagnostic,
  title = {{LLM-Agnostic Semantic Representation Attack}},
  author = {Jiawei Lian and Jianhong Pan and Lefan Wang and Yi Wang and Tairan Huang and Shaohui Mei and Lap-Pui Chau},
  year = {2026},
  month = may,
  eprint = {2605.08898},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.08898}
}