July 2026Unreviewed
RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
Benyamin Tafreshian, Prathamesh Dhake
Abstract
Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls. Despite these controls, preventing prompt-based policy evasion remains challenging, and adversaries continue to "jailbreak" LLMs by crafting prompts that circumvent implemented safety mechanisms. Prior work has established cipher-mediated interaction, code-embedded decryption, prompt decomposition and reconstru
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{tafreshian2026rogueprompt,
title = {{RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation}},
author = {Benyamin Tafreshian and Prathamesh Dhake},
year = {2026},
month = jul,
eprint = {2607.27373},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.27373}
}