April 2026Unreviewed
SafeDream: Safety World Model for Proactive Early Jailbreak Detection
Bo Yan, Weikai Lin, Yada Zhu, Song Wang
Abstract
Multi-turn jailbreak attacks progressively erode LLM safety alignment across seemingly innocuous conversation turns, achieving success rates exceeding 90% against state-of-the-art models. Existing alignment-based and guardrail methods suffer from three key limitations: they require costly weight modification, evaluate each turn independently without modeling cumulative safety erosion, and detect attacks only after harmful content has been generated. To address these limitations, we first formula
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{yan2026safedream,
title = {{SafeDream: Safety World Model for Proactive Early Jailbreak Detection}},
author = {Bo Yan and Weikai Lin and Yada Zhu and Song Wang},
year = {2026},
month = apr,
eprint = {2604.16824},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.16824}
}