July 2026Unreviewed
Do LLMs Know Their Vulnerable Scenarios?
Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
Abstract
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relation
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{peng2026do,
title = {{Do LLMs Know Their Vulnerable Scenarios?}},
author = {Ziheng Peng and Huiqi Deng and Haoran Jing and Xuankun Rong and Jiahui Han and Xiting Wang and Na Zou and Xia Hu},
year = {2026},
month = jul,
eprint = {2607.23496},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.23496}
}