← Back to search
paper llmsec-2026-00075
LLM-Agnostic Semantic Representation Attack
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Tairan Huang, Shaohui Mei, Lap-Pui Chau
2026-05
Abstract
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting adversarial prompts. Predominant token-level optimization methods primarily rely on optimizing for exact affirmative templates (e.g., ``\textit{Sure, here is...}''). However, these paradigms frequently encounter bottlenecks such as suboptimal convergence, compromised prompt naturalness, and poor cross-model generalization. To addres
Categories
Cite This Resource
@article{llmsec202600075,
title = {LLM-Agnostic Semantic Representation Attack},
author = {Jiawei Lian and Jianhong Pan and Lefan Wang and Yi Wang and Tairan Huang and Shaohui Mei and Lap-Pui Chau},
year = {2026},
url = {https://arxiv.org/abs/2605.08898},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.08898