August 2025Unreviewed
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
Wenpeng Xing, Mohan Li, Chunqiang Hu, Haitao Zhang, Bo Lin, Meng Han
arXiv.org
Abstract
While Large Language Models (LLMs) have achieved remarkable progress, they remain vulnerable to jailbreak attacks. Existing methods, primarily relying on discrete input optimization (e.g., GCG), often suffer from high computational costs and generate high-perplexity prompts that are easily blocked by simple filters. To overcome these limitations, we propose Latent Fusion Jailbreak (LFJ), a stealthy white-box attack that operates in the continuous latent space. Unlike previous approaches, LFJ con
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{xing2025latent,
title = {{Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs}},
author = {Wenpeng Xing and Mohan Li and Chunqiang Hu and Haitao Zhang and Bo Lin and Meng Han},
year = {2025},
month = aug,
eprint = {2508.10029},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2508.10029},
url = {https://www.semanticscholar.org/paper/3e9ca9bb9cd160d045ee046e026bc6efd8b762e0}
}