Skip to content
Search
paperMay 2026Unreviewed

Adversarial Reframing: A Framework for Targeted Generation in Language Models

Shahnewaz Karim Sakib, Swati Kar, Anindya Bijoy Das

Abstract

Large Language Models (LLMs) are widely deployed in diverse real-world settings, yet remain vulnerable to jailbreaking, where prompt-based attacks bypass safety filters. We present THREAT (Targeted Harmful generation via Reframing and Exploitation of Adversarial Tactics), a reasoning-driven framework that coordinates multiple LLMs in an iterative search loop to find textual jailbreak prompts. We formulate prompt discovery as a nonconvex optimization problem and provide an efficient solution that

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{sakib2026adversarial,
  title = {{Adversarial Reframing: A Framework for Targeted Generation in Language Models}},
  author = {Shahnewaz Karim Sakib and Swati Kar and Anindya Bijoy Das},
  year = {2026},
  month = may,
  eprint = {2605.21674},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.21674}
}