Skip to content
Search
paperSeptember 2026Unreviewed

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

Thomas Rivasseau

Abstract

Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak and have previously been demonstrated against the fine-tuning APIs of commercial models. In these attacks, target models are trained on a corpus of encrypted harmful questions and re

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{rivasseau2026arbitrary,
  title = {{Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning}},
  author = {Thomas Rivasseau},
  year = {2026},
  month = sep,
  eprint = {2609.09553},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.09553}
}