May 2026Unreviewed
EVA: Editing for Versatile Alignment against Jailbreaks
Yi Wang, Hongye Qiu, Yue Xu, Sibei Yang, Zhan Qin, Minlie Huang, Wenjie Wang
Abstract
Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks, where adversaries exploit textual or visual triggers to bypass safety guardrails. Recent defenses typically rely on safety fine-tuning or external filters to reduce the model's likelihood of producing harmful content. While effective to some extent, these methods often incur significant computational overheads and suffer from the safety utility t
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{wang2026eva,
title = {{EVA: Editing for Versatile Alignment against Jailbreaks}},
author = {Yi Wang and Hongye Qiu and Yue Xu and Sibei Yang and Zhan Qin and Minlie Huang and Wenjie Wang},
year = {2026},
month = may,
eprint = {2605.14750},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.14750}
}