September 2026Unreviewed
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
Thomas Rivasseau
Abstract
Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arbitrary cipher, or covert communication, attacks are one such type of jailbreak and have previously been demonstrated against the fine-tuning APIs of commercial models. In these attacks, target models are trained on a corpus of encrypted harmful questions and re
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{rivasseau2026arbitrary,
title = {{Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning}},
author = {Thomas Rivasseau},
year = {2026},
month = sep,
eprint = {2609.09553},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.09553}
}