Skip to content

Red Teaming

AI red team methodology, automation, and frameworks

Resources
94
Page
2/2

Newest first · 11 reviewed on this page

Search instead
paper2026Stout in Computer Science and Technology StudiesUnreviewed

Continual Red-Teaming and Guardrail Distillation for Tool-Using LLM Agents: Prompt-Injection Resistance with Utility Preservation

Wesley Gao

Tool-using language-model agents can convert indirect prompt injection into consequential actions, making guardrail quality a joint security, utility, and efficiency problem. This study evaluates a ReAct-style control, native tool filtering, deterministic self-verification, a…

paper2025arXiv.orgUnreviewed

Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs

Chetan Pathade

Large Language Models (LLMs) are increasingly integrated into consumer and enterprise applications. Despite their capabilities, they remain susceptible to adversarial attacks such as prompt injection and jailbreaks that override alignment safeguards. This paper provides a…

paper2025North American Chapter of the Association for Computational LinguisticsUnreviewed

RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

Shiyue Zhang, M. Dredze, AI Bloomberg +77

Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. However, despite the widespread use of the Retrieval-Augmented Generation (RAG) framework, AI safety work focuses on standard LLMs, which means we know little…

paper2025AAAI Conference on Artificial IntelligenceUnreviewed

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Ian R. McKenzie, O. Hollinsworth, Tom Tseng +5

Frontier AI developers are relying on layers of safeguards to protect against catastrophic misuse of AI systems. Anthropic guards their latest Claude 4 Opus model using one such defense pipeline, and other frontier developers including Google DeepMind and OpenAI pledge to soon…

paper2025Conference of the European Chapter of the Association for Computational LinguisticsUnreviewed

SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning

Kai Zhou, Ahmed Elgohary, S. M. Iftekhar +80

The ability of LLM agents to plan and invoke tools exposes them to new safety risks, making a comprehensive red-teaming system crucial for discovering vulnerabilities and ensuring their safe deployment. We present SIRAJ: a generic red-teaming framework for arbitrary black-box…

paper2024AAAI/ACM Conference on AI, Ethics, and SocietyUnreviewed

Red-Teaming for Generative AI: Silver Bullet or Security Theater?

Michael Feffer, Anusha Sinha, Zachary Chase Lipton +1

In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying and mitigating these risks. However,…

Red Teaming151 cit.
paper2024International Conference on Machine LearningUnreviewed

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Xiangming Gu, Xiaosen Zheng, Tianyu Pang +5

A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned…

paper2024International Conference on Machine LearningUnreviewed

A Safe Harbor for AI Evaluation and Red Teaming

Shayne Longpre, Sayash Kapoor, Kevin Klyman +20

Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by prominent AI companies to deter model misuse have disincentives on good faith safety evaluations. This…

paper2023Conference on Empirical Methods in Natural Language ProcessingUnreviewed

AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications

Bhaktipriya Radharapu, Kevin Robinson, L. Aroyo +1

Adversarial testing of large language models (LLMs) is crucial for their safe and responsible deployment. We introduce a novel approach for automated generation of adversarial evaluation datasets to test the safety of LLM generations on new downstream applications. We call it…

paper2023Conference on Empirical Methods in Natural Language ProcessingUnreviewed

ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Alex Mei, Sharon Levy, W. Wang

As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustness evaluations must comprehensively encapsulate the various settings in which a user may invoke an…