Skip to content
Search
paperJuly 2026Unreviewed

RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation

Benyamin Tafreshian, Prathamesh Dhake

Abstract

Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls. Despite these controls, preventing prompt-based policy evasion remains challenging, and adversaries continue to "jailbreak" LLMs by crafting prompts that circumvent implemented safety mechanisms. Prior work has established cipher-mediated interaction, code-embedded decryption, prompt decomposition and reconstru

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{tafreshian2026rogueprompt,
  title = {{RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation}},
  author = {Benyamin Tafreshian and Prathamesh Dhake},
  year = {2026},
  month = jul,
  eprint = {2607.27373},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.27373}
}