Skip to content

Red Teaming

AI red team methodology, automation, and frameworks

Resources
94
Page
1/2

Newest first

Search instead
paper20262026 International Conference on Intelligent Multimedia, Networking, and Security (IMNS)Unreviewed

Whispers of Wealth: A Systematic Red-Teaming Study of the Agent Payments Protocol (AP2)

Tanusree Debi, Wentian Zhu

Large language model (LLM)-based agents are increasingly used to automate financial transactions, but their reliance on contextual reasoning introduces new security risks. The Agent Payments Protocol (AP2) secures agent-mediated purchases through cryptographically signed…

paper2026arXiv.orgUnreviewed

Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming

I. Steenstra, Paola Pedrelli, Weiyan Shi +2

Large Language Models (LLMs) are increasingly utilized for mental health support; however, current safety benchmarks often fail to detect the complex, longitudinal risks inherent in therapeutic dialogue. We introduce an evaluation framework that pairs AI psychotherapists with…

paper2026Unreviewed

FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption

Yanting Wang, Chenlong Yin, Ying Chen +1

Long-context large language models (LLMs)-for example, Gemini-3.1-Pro and Qwen-3.5-are widely used to empower many real-world applications, such as retrieval-augmented generation, autonomous agents, and AI assistants. However, security remains a major concern for their…

paper2026Unreviewed

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

Nicholas Saban

Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster on retired models and on the most-vulnerable model in each paper's panel. We ask whether those techniques, reproduced as…

paper20262026 8th International Conference on Software Engineering and Computer Science (CSECS)Unreviewed

Double-Gaming: Jailbreak Attacks against LLM based on Red-Blue Team Game Theory

Chenlu Ma, Huairui Zhao, G. Nie +3

Despite existing security alignment mechanisms, Large Language Models (LLMs) remain vulnerable to jailbreak attacks under static defenses. To address this, we propose a novel jailbreak attack and defense optimization framework based on Red-Blue Team dynamic game theory. This…

paper2026Unreviewed

RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

Yarin Yerushalmi Levi, Roy Betser, Amit Giloni +5

Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains,…

paper2026Unreviewed

Do LLMs Know Their Vulnerable Scenarios?

Ziheng Peng, Huiqi Deng, Haoran Jing +5

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why…

paper2026Unreviewed

GPT-Red: Automated Red Teaming via Self-Play at Scale

Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15

We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially…

paper2026Unreviewed

Generating Attacks for LLMs with GFlowNets

Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali +1

The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this escalating trend has introduced significant security vulnerabilities, necessitating the identification and…

paper20262026 IEEE 2nd International Conference on Secure IoT, Assured and Trusted Computing (SATC)Unreviewed

Benchmarking the Effectiveness of AI-Driven Red Teaming Across Safety-Aligned Language Models

Joyce Malicha, Kamrul Hasan

Large language models (LLMs) have advanced rapidly, yet even safety-aligned models remain vulnerable to adversarial prompts that bypass safeguards and induce harmful outputs. Conventional red teaming methods, including static testing and gradient-based attacks, are limited by…

paper2026Unreviewed

Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu +3

Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize these failures through a…

paper2026Unreviewed

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

Feitong Qiao, Liren Peng, Shiming Ren +7

Frontier language models that refuse harmful single-turn prompts often comply when the same intent is reached gradually over many turns, making multi-turn attacks one of the least understood failure modes of large language models. Most automated red-teaming methods treat this as…

paper2026Unreviewed

SIR: Self-improving Red-teaming for Compute Use Agents

Chen Xiong, Zhiyuan He, Pin-Yu Chen +2

Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate everyday digital tasks. Because they can be exposed to untrusted content while…