Skip to content
Search
paperJuly 2026Unreviewed

Do LLMs Know Their Vulnerable Scenarios?

Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu

Abstract

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relation

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{peng2026do,
  title = {{Do LLMs Know Their Vulnerable Scenarios?}},
  author = {Ziheng Peng and Huiqi Deng and Haoran Jing and Xuankun Rong and Jiahui Han and Xiting Wang and Na Zou and Xia Hu},
  year = {2026},
  month = jul,
  eprint = {2607.23496},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.23496}
}