May 2026Unreviewed
Adversarial Reframing: A Framework for Targeted Generation in Language Models
Shahnewaz Karim Sakib, Swati Kar, Anindya Bijoy Das
Abstract
Large Language Models (LLMs) are widely deployed in diverse real-world settings, yet remain vulnerable to jailbreaking, where prompt-based attacks bypass safety filters. We present THREAT (Targeted Harmful generation via Reframing and Exploitation of Adversarial Tactics), a reasoning-driven framework that coordinates multiple LLMs in an iterative search loop to find textual jailbreak prompts. We formulate prompt discovery as a nonconvex optimization problem and provide an efficient solution that
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{sakib2026adversarial,
title = {{Adversarial Reframing: A Framework for Targeted Generation in Language Models}},
author = {Shahnewaz Karim Sakib and Swati Kar and Anindya Bijoy Das},
year = {2026},
month = may,
eprint = {2605.21674},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.21674}
}