Skip to content
Search
paperJuly 2026Unreviewed

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra, Victoria Krakovna

Abstract

Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasion-based jailbreaks, where natural-language arguments override model constraints. We stress-test whether this vulnerability extends to monitoring LLMs: can an adversarial agent persuade its CoT monitor to approve proposed

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{za2026persuasion,
  title = {{Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring}},
  author = {Jennifer Za and Julija Bainiaksina and Nikita Ostrovsky and Tanush Chopra and Victoria Krakovna},
  year = {2026},
  month = jul,
  eprint = {2607.08066},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.08066}
}