Skip to content
Search
paperMarch 2026Unreviewed

Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning

Bilgehan Sel, Xuanli He, Alwin Peng, Ming Jin, Jerry Wei

arXiv.org

Abstract

Fine-tuning APIs offered by major AI providers create new attack surfaces where adversaries can bypass safety measures through targeted fine-tuning. We introduce Trojan-Speak, an adversarial fine-tuning method that bypasses Anthropic's Constitutional Classifiers. Our approach uses curriculum learning combined with GRPO-based hybrid reinforcement learning to teach models a communication protocol that evades LLM-based content classification. Crucially, while prior adversarial fine-tuning approache

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{sel2026trojanspeak,
  title = {{Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning}},
  author = {Bilgehan Sel and Xuanli He and Alwin Peng and Ming Jin and Jerry Wei},
  year = {2026},
  month = mar,
  eprint = {2603.29038},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2603.29038},
  url = {https://www.semanticscholar.org/paper/97dad0cd24d00035e5ae90c101de5072c2d1174d}
}