Skip to content

Jailbreaking

Guardrail bypass and alignment subversion techniques

Resources
273
Page
5/6

Newest first · 4 reviewed on this page

Search instead
paper2026CybersecurityUnreviewed

RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling

Michael Tchuindjang, Nathan Duran, Phil Legg +1

Existing multi-turn jailbreak evaluation methods increasingly rely on large language models (LLMs) as automated judges to reduce the cost and scalability limitations of human assessment. However, recent studies show that LLM-based evaluators can diverge from human judgments…

JailbreakingOpen access
paper2026ElectronicsUnreviewed

Securing the Prompt Pipeline: A Systematic Review of Defense Mechanisms Against Prompt-Based Attacks in LLM Agents

Sana Mourad, E. E. Abdallah, Mohammad Ababneh

Current language model deployments face growing security challenges from prompt-based attacks, including jailbreaks, direct and indirect prompt injection, and instruction hijacking, which often evade traditional rule-based safeguards. As these models are increasingly integrated…

paper2026Unreviewed

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

Jiawei Liu, Jiacheng Guo, Tian Zhang +4

Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks,…

paper2026Unreviewed

Measuring Semantic Abstractness of SAE Features via Nonlocality

Chuqi Lin, S. Sondhi, Xiao-Liang Qi

Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must…

paper2026ElectronicsUnreviewed

A Privacy-Preserving Middleware Architecture for Detecting Prompt Injection and Sensitive Data Exposure in Large-Language-Model Interactions

Adam Ait Hsine, A. Arabo

The deployment of large language models (LLMs) in real-world applications introduces a compounding security problem: detecting adversarial inputs such as prompt injection and jailbreak-driven data leakage while simultaneously preventing the detection mechanism itself from…

paper2026Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2Unreviewed

The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models

Lu Lin, Jinghui Chen, Ting Wang +4

Large language models (LLMs) are increasingly embedded as core components of data-centric systems, supporting analytical decision making, and automated reasoning over large-scale, heterogeneous datasets. Yet their deployment in open-world environments raises fundamental…

paper2026Unreviewed

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt +1

Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense…

paper2025arXiv.orgUnreviewed

Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs

Chetan Pathade

Large Language Models (LLMs) are increasingly integrated into consumer and enterprise applications. Despite their capabilities, they remain susceptible to adversarial attacks such as prompt injection and jailbreaks that override alignment safeguards. This paper provides a…

paper2025Conference on Empirical Methods in Natural Language ProcessingUnreviewed

Foot-In-The-Door: A Multi-turn Jailbreak for LLMs

Zixuan Weng, Xiaolong Jin, Jinyuan Jia +1

Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs. Inspired by psychological foot-in-the-door…

paper2025Conference on Empirical Methods in Natural Language ProcessingUnreviewed

AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender

Weixiang Zhao, Jiahe Guo, Yulin Hu +8

Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false…

paper2025Unreviewed

Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems

William Hackett, Lewis Birch, Stefan Trawicki +2

Large Language Models (LLMs) guardrail systems are designed to protect against prompt injection and jailbreak attacks. However, they remain vulnerable to evasion techniques. We demonstrate two approaches for bypassing LLM prompt injection and jailbreak detection systems via…

paper2025Conference on Empirical Methods in Natural Language ProcessingUnreviewed

Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation

Daniel Schwartz, Dmitriy Bespalov, Zhe Wang +2

As large language models (LLMs) become increasingly prevalent, ensuring their robustness against adversarial misuse is crucial. This paper introduces the GAP (Graph of Attacks with Pruning) framework, an advanced approach for generating stealthy jailbreak prompts to evaluate and…

paper2025Annual Meeting of the Association for Computational LinguisticsUnreviewed

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

Xiaohu Li, Yunfeng Ning, Zepeng Bao +3

Security alignment enables the Large Language Model (LLM) to gain the protection against malicious queries, but various jailbreak attack methods reveal the vulnerability of this security mechanism. Previous studies have isolated LLM jailbreak attacks and defenses. We analyze the…

paper2025arXiv.orgUnreviewed

Proactive defense against LLM Jailbreak

Weiliang Zhao, Jinjun Peng, Daniel Ben-Levi +2

The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including multi-turn jailbreaks that iteratively search for successful queries. Current defenses, which are…

paper2025AAAI Conference on Artificial IntelligenceUnreviewed

AlignTree: Efficient Defense Against LLM Jailbreak Attacks

Gil Goren, Shahar Katz, Lior Wolf

Large Language Models (LLMs) are vulnerable to adversarial attacks that bypass safety guidelines and generate harmful content. Mitigating these vulnerabilities requires defense mechanisms that are both robust and computationally efficient. However, existing approaches either…

paper2025arXiv.orgUnreviewed

Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization

Xurui Li, Kaisong Song, Rui Zhu +2

Large Language Models (LLMs) have developed rapidly in web services, delivering unprecedented capabilities while amplifying societal risks. Existing works tend to focus on either isolated jailbreak attacks or static defenses, neglecting the dynamic interplay between evolving…

paper20252025 International Conference on Emerging Information Technology and Engineering Solutions (EITES)Unreviewed

Breaking the Shield: Adversarial Jailbreak Attacks and Defense Mechanisms in Large Language Models

Shreya Dubey, H. Lamkuche

Large Language Models (LLMs) are increasingly used across diverse applications, raising security concerns, especially regarding jail-break attacks that bypass safety measures. This paper presents a systematic security evaluation of LLMs, exploring the effectiveness of jail-break…