July 2026Unreviewed
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra, Victoria Krakovna
Abstract
Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasion-based jailbreaks, where natural-language arguments override model constraints. We stress-test whether this vulnerability extends to monitoring LLMs: can an adversarial agent persuade its CoT monitor to approve proposed
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{za2026persuasion,
title = {{Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring}},
author = {Jennifer Za and Julija Bainiaksina and Nikita Ostrovsky and Tanush Chopra and Victoria Krakovna},
year = {2026},
month = jul,
eprint = {2607.08066},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.08066}
}