Skip to content
Search
paperAugust 2025Unreviewed

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

Wenpeng Xing, Mohan Li, Chunqiang Hu, Haitao Zhang, Bo Lin, Meng Han

arXiv.org

Abstract

While Large Language Models (LLMs) have achieved remarkable progress, they remain vulnerable to jailbreak attacks. Existing methods, primarily relying on discrete input optimization (e.g., GCG), often suffer from high computational costs and generate high-perplexity prompts that are easily blocked by simple filters. To overcome these limitations, we propose Latent Fusion Jailbreak (LFJ), a stealthy white-box attack that operates in the continuous latent space. Unlike previous approaches, LFJ con

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{xing2025latent,
  title = {{Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs}},
  author = {Wenpeng Xing and Mohan Li and Chunqiang Hu and Haitao Zhang and Bo Lin and Meng Han},
  year = {2025},
  month = aug,
  eprint = {2508.10029},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2508.10029},
  url = {https://www.semanticscholar.org/paper/3e9ca9bb9cd160d045ee046e026bc6efd8b762e0}
}