← Back to search
paper llmsec-2026-00075

LLM-Agnostic Semantic Representation Attack

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Tairan Huang, Shaohui Mei, Lap-Pui Chau

2026-05

Abstract

Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting adversarial prompts. Predominant token-level optimization methods primarily rely on optimizing for exact affirmative templates (e.g., ``\textit{Sure, here is...}''). However, these paradigms frequently encounter bottlenecks such as suboptimal convergence, compromised prompt naturalness, and poor cross-model generalization. To addres

Categories

Cite This Resource

@article{llmsec202600075,
  title = {LLM-Agnostic Semantic Representation Attack},
  author = {Jiawei Lian and Jianhong Pan and Lefan Wang and Yi Wang and Tairan Huang and Shaohui Mei and Lap-Pui Chau},
  year = {2026},
  url = {https://arxiv.org/abs/2605.08898},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.08898