Frameworks
Research, mapped to the standards
Find the papers and tools behind each OWASP risk, MITRE ATLAS technique, NIST AI RMF function and ISO/IEC 42001 clause. Mappings on unreviewed entries are suggested from their categories.
OWASP Top 10 for LLM Applications
v2025 · 725 mapped resources
Prompt Injection
Manipulating LLMs through crafted inputs
- Universal and Transferable Adversarial Attacks on Aligned Language Models2023
- Jailbroken: How Does LLM Safety Training Fail?2024
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection2023
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly2024
Show 8 moreShow fewer
- Do Anything Now: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models2023
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models2023
- Prompt Injection Attack Against LLM-Integrated Applications2024
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations2023
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks2024
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models2024
- Jailbreaking Black Box Large Language Models in Twenty Queries2025
- Visual Adversarial Examples Jailbreak Aligned Large Language Models2024
Most cited shown; 522 more are mapped here.
Sensitive Information Disclosure
Unintended revelation of confidential data
- Extracting Training Data from Large Language Models2021
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly2024
- Scalable Extraction of Training Data from (Production) Language Models2023
- Multi-step Jailbreaking Privacy Attacks on ChatGPT2023
Show 8 moreShow fewer
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory2024
- A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language Models2024
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks2024
- Garak: A Framework for Security Probing Large Language Models2024
- Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models2024
- Mitigating adversarial manipulation in LLMs: a prompt-based approach to counter Jailbreak attacks (Prompt-G)2024
- SecureGov-Agent: A Governance-Centric Multi-Agent Framework for Privacy-Preserving and Attack-Resilient LLM Agents2025
- Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models2026
Most cited shown; 54 more are mapped here.
Supply Chain
Compromised components in LLM supply chain
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly2024
- Poisoning Web-Scale Training Datasets is Practical2024
- Stealing Part of a Production Language Model2024
- A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language Models2024
Show 8 moreShow fewer
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models2024
- GPT in Sheep's Clothing: The Risk of Customized GPTs2024
- Formal Analysis and Supply Chain Security for Agentic AI Skills2026
- Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation2025
- Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain2026
- Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning2026
- Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection2026
- GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines2026
Most cited shown; 10 more are mapped here.
Data and Model Poisoning
Tampering with training data or models
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2024
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly2024
- Poisoning Web-Scale Training Datasets is Practical2024
- Poisoning Language Models During Instruction Tuning2023
Show 8 moreShow fewer
- LoRA Fine-Tuning Efficiently Undoes Safety Training in Llama 2-Chat2023
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications2024
- A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language Models2024
- PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models2024
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models2024
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models2023
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models2024
- SecureGov-Agent: A Governance-Centric Multi-Agent Framework for Privacy-Preserving and Attack-Resilient LLM Agents2025
Most cited shown; 157 more are mapped here.
Improper Output Handling
Insufficient validation of LLM outputs
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations2023
- NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails2023
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models2024
- A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language Models2024
Show 7 moreShow fewer
- From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?2024
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs2024
- OWASP Top 10 for Large Language Model Applications2025
- OWASP AI Security and Privacy Guide2024
- OWASP LLM AI Security & Governance Checklist2024
- Guardrails AI: Input/Output Guards for LLM Applications2024
- LLM Guard: Security Toolkit for LLM Interactions2024
Excessive Agency
Granting LLMs too much autonomy or access
- Toolformer: Language Models Can Teach Themselves to Use Tools2023
- LLM Agents Can Autonomously Hack Websites2024
- NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails2023
- LLM Agents Can Autonomously Exploit One-day Vulnerabilities2024
Show 6 moreShow fewer
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents2024
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents2024
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents2024
- OWASP Top 10 for Large Language Model Applications2025
- OWASP AI Security and Privacy Guide2024
- OWASP LLM AI Security & Governance Checklist2024
System Prompt Leakage
Exposure of system-level instructions
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions2024
- Prompt Stealing Attacks Against Text-to-Image Generation Models2024
- GPT in Sheep's Clothing: The Risk of Customized GPTs2024
- OWASP Top 10 for Large Language Model Applications2025
Show 1 moreShow fewer
Vector and Embedding Weaknesses
Vulnerabilities in vector stores and embeddings
Misinformation
Generation of false or misleading content
Unbounded Consumption
Uncontrolled resource usage by LLMs
OWASP Top 10 for Agentic Applications
v2026 · 70 mapped resources
Agent Goal Hijack
Manipulating an agent's goals, plans or decision path, e.g. through injected instructions
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents2024
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents2024
- The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies2024
- Dissecting Adversarial Robustness of Multimodal LM Agents2025
Tool Misuse & Exploitation
Agents using legitimate tools in unsafe or unintended ways
- ReAct: Synergizing Reasoning and Acting in Language Models2023
- Toolformer: Language Models Can Teach Themselves to Use Tools2023
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering2024
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents2024
Show 8 moreShow fewer
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents2024
- The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies2024
- Dissecting Adversarial Robustness of Multimodal LM Agents2025
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security2026
- SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment2026
- Securing LLM-based agents against cyberattacks: a comprehensive survey on attack techniques and defense strategies2026
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks2026
- ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection2026
Most cited shown; 33 more are mapped here.
Agent Identity & Privilege Abuse
Exploiting delegated identities, credentials and excessive privileges
- ReAct: Synergizing Reasoning and Acting in Language Models2023
- LLM Agents Can Autonomously Hack Websites2024
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering2024
- LLM Agents Can Autonomously Exploit One-day Vulnerabilities2024
Show 2 moreShow fewer
Agentic Supply Chain Compromise
Compromised tools, plugins, models, prompts or agents pulled in at runtime
Unexpected Code Execution
Agent-generated or agent-triggered code execution leading to compromise
Memory & Context Poisoning
Corrupting stored memory, RAG data or context that shapes future behavior
- Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents2026
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning2026
- Memory poisoning and secure multi-agent systems2026
- Hidden in Memory: Sleeper Memory Poisoning in LLM Agents2026
Show 8 moreShow fewer
- Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction2026
- Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents2026
- The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems2026
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents2026
- MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents2026
- Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees2026
- Forensic Trajectory Signatures for Agent Memory Poisoning Detection2026
- When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents2026
Most cited shown; 10 more are mapped here.
Insecure Inter-Agent Communication
Unauthenticated or unencrypted agent-to-agent messages enabling spoofing and manipulation
Cascading Failures
Faults amplifying across coupled agents and systems
Human-Agent Trust Exploitation
Exploiting users' trust in agents, e.g. with fabricated justifications
Rogue Agents
Agents deviating from intended behavior without active attacker control
- Voyager: An Open-Ended Embodied Agent with Large Language Models2023
- LLM Agents Can Autonomously Hack Websites2024
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering2024
- LLM Agents Can Autonomously Exploit One-day Vulnerabilities2024
Show 1 moreShow fewer
MITRE ATLAS
v5.6.0 · 768 mapped resources
AI Supply Chain Compromise
Adversaries may gain initial access to a system by compromising the unique portions of the AI supply chain.
- Formal Analysis and Supply Chain Security for Agentic AI Skills2026
- Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation2025
- Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain2026
- Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning2026
Show 8 moreShow fewer
- Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection2026
- GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines2026
- MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills2026
- Trustworthy AI LLM Scalability Risk Index (LSRI): A Cybersecurity Framework Assessing Agentic-AI Security & Software Model Supply Chain Safety Boosting AI-Generated Malware Defense & Explainability Mitigating Emerging Risks of Generative AI2026
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering2026
- VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities2026
- When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination2026
- ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization2026
Most cited shown; 2 more are mapped here.
Evade AI Model
Adversaries can Craft Adversarial Data that prevents an AI model from correctly identifying the contents of the data or Generate Deepfakes that fools an AI m...
Manipulate AI Model
Adversaries may directly manipulate an AI model to change its behavior or introduce malicious code.
Publish Poisoned Datasets
Adversaries may Poison Training Data and publish it to a public location.
Poison Training Data
Adversaries may attempt to poison datasets used by an AI model by modifying the underlying data or its labels.
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training2024
- Poisoning Web-Scale Training Datasets is Practical2024
- Poisoning Language Models During Instruction Tuning2023
- PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models2024
Show 8 moreShow fewer
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models2024
- SecureGov-Agent: A Governance-Centric Multi-Agent Framework for Privacy-Preserving and Attack-Resilient LLM Agents2025
- ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications2025
- RoLLMRec: a robust LLM-based recommender system for defending against shilling and prompt injection attacks2026
- HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking2026
- Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning2026
- CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents2026
- MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs2026
Most cited shown; 149 more are mapped here.
Exfiltration via AI Inference API
Adversaries may exfiltrate private information via AI Model Inference API Access.
Infer Training Data Membership
Adversaries may infer the membership of a data sample or global characteristics of the data in its training set, which raises privacy concerns.
- Mitigating adversarial manipulation in LLMs: a prompt-based approach to counter Jailbreak attacks (Prompt-G)2024
- SecureGov-Agent: A Governance-Centric Multi-Agent Framework for Privacy-Preserving and Attack-Resilient LLM Agents2025
- Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models2026
- Membership Inference Attacks Against Video Large Language Models2026
Show 8 moreShow fewer
- Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks2026
- Toward Efficient Membership Inference Attacks against Federated Large Language Models: A Projection Residual Approach2026
- Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs2026
- Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications2026
- Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System2026
- An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments2026
- AI Security Research Should Better Incentivize Defense Research2026
- A Unified Evaluation Framework for Utility and Privacy Risks of LLM-Generated Synthetic Text Data2026
Most cited shown; 41 more are mapped here.
Extract AI Model
Adversaries may extract a functional copy of a private model.
- Prompt Stealing Attacks Against Text-to-Image Generation Models2024
- Stealing Part of a Production Language Model2024
- Attack and defense techniques in large language models: A survey and new perspectives2025
- Adaptive Probe-based Steering for Robust LLM Jailbreaking2026
Show 8 moreShow fewer
- An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic2026
- Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot2026
- JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols2026
- AI-Assisted Extraction of Follow-up Observations from GCN Circulars in Astro-COLIBRI2026
- Towards Automated Domain Model Extraction from Source Code using Heuristics and Open-Source LLMs2026
- Vision-Language Semantics Guided Model Extraction Attack via Long-Horizon Contrastive Prompt Learning2026
- Adversarial Machine Learning: Attack Vectors, Defences, and Robustness2026
- MITRE ATLAS: Adversarial Threat Landscape for AI Systems2024
Exfiltration via Cyber Means
Adversaries may exfiltrate AI artifacts or other information relevant to their goals via traditional cyber means.
Denial of AI Service
Adversaries may target AI-enabled systems with a flood of requests for the purpose of degrading or shutting down the service.
Cost Harvesting
Adversaries may deliberately drive a victim's AI services beyond normal operating capacity with the intent of increasing the cost of services.
AI Model Inference API Access
Adversaries may gain access to a model via legitimate access to the inference API.
Verify Attack
Adversaries can verify the efficacy of their attack via an inference API or access to an offline copy of the target model.
Craft Adversarial Data
Adversarial data are inputs to an AI model that have been modified such that they cause the adversary's desired effect in the target model.
- Universal and Transferable Adversarial Attacks on Aligned Language Models2023
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models2023
- Visual Adversarial Examples Jailbreak Aligned Large Language Models2024
- Adversarial Attacks and Defenses in Large Language Models: Old and New Threats2024
Show 8 moreShow fewer
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models2024
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs2025
- Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judge2025
- Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems2025
- STACK: Adversarial Attacks on LLM Safeguard Pipelines2025
- CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations2025
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness2026
- Proactive defense against LLM Jailbreak2025
Most cited shown; 60 more are mapped here.
Full AI Model Access
Adversaries may gain full "white-box" access to an AI model.
AI-Enabled Product or Service
Adversaries may use a product or service that uses artificial intelligence under the hood to gain access to the underlying AI model.
LLM Prompt Injection
An adversary may craft malicious prompts as inputs to an LLM that cause the LLM to act in unintended ways.
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection2023
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly2024
- Prompt Injection Attack Against LLM-Integrated Applications2024
- Ignore This Title and HackAPrompt: Exposing Systemic Weaknesses of LLMs through a Global Scale Prompt Hacking Competition2023
Show 8 moreShow fewer
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions2024
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game2024
- Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications2024
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents2024
- From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?2024
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models2024
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses2024
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents2025
Most cited shown; 264 more are mapped here.
AI Agent Tool Invocation
Adversaries may use their access to an AI agent to invoke tools the agent has access to.
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents2024
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security2026
- SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment2026
- Securing LLM-based agents against cyberattacks: a comprehensive survey on attack techniques and defense strategies2026
Show 8 moreShow fewer
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks2026
- ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection2026
- When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents2026
- Beyond the Black Box: Interpretability of Agentic AI Tool Use2026
- Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use2026
- AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use2026
- Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models2026
- AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum2026
Most cited shown; 24 more are mapped here.
LLM Jailbreak
Adversaries may induce a large language model (LLM) to ignore, circumvent, or override its safety/alignment behaviors and/or guardrails to elicit outputs the...
- Universal and Transferable Adversarial Attacks on Aligned Language Models2023
- Jailbroken: How Does LLM Safety Training Fail?2024
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly2024
- Do Anything Now: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models2023
Show 8 moreShow fewer
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models2023
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks2024
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models2024
- Jailbreaking Black Box Large Language Models in Twenty Queries2025
- Ignore This Title and HackAPrompt: Exposing Systemic Weaknesses of LLMs through a Global Scale Prompt Hacking Competition2023
- Tree of Attacks: Jailbreaking Black-Box LLMs with Auto-Generated Subtrees2024
- Multi-step Jailbreaking Privacy Attacks on ChatGPT2023
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher2024
Most cited shown; 259 more are mapped here.
Extract LLM System Prompt
Adversaries may attempt to extract a large language model's (LLM) system prompt.
LLM Data Leakage
Adversaries may craft prompts that induce the LLM to leak sensitive information.
Publish Poisoned Models
Adversaries may publish a poisoned model to a public location such as a model registry or code repository.
RAG Poisoning
Adversaries may inject malicious content into data indexed by a retrieval augmented generation (RAG) system to contaminate a future thread through RAG-based...
AI Agent Context Poisoning
Adversaries may attempt to manipulate the context used by an AI agent's large language model (LLM) to influence the responses it generates or actions it takes.
- Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents2026
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning2026
- Memory poisoning and secure multi-agent systems2026
- Hidden in Memory: Sleeper Memory Poisoning in LLM Agents2026
Show 8 moreShow fewer
- Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction2026
- Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents2026
- The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems2026
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents2026
- MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents2026
- Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees2026
- Forensic Trajectory Signatures for Agent Memory Poisoning Detection2026
- When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents2026
Most cited shown; 9 more are mapped here.
NIST AI Risk Management Framework
v1.0 · 111 mapped resources
Govern
Cultivate a culture of risk management
- Constitutional AI: Harmlessness from AI Feedback2022
- Identifying and Mitigating the Security Risks of Generative AI2023
- On the Societal Impact of Open Foundation Models2024
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory2024
Show 8 moreShow fewer
- Machine Unlearning in Generative AI: A Survey2024
- OWASP AI Security and Privacy Guide2024
- OWASP LLM AI Security & Governance Checklist2024
- An Architectural Risk Analysis of Large Language Models: Applied Machine Learning Security2024
- Anthropic's Responsible Scaling Policy2024
- Generative AI Security: Theories and Practices2024
- NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0)2023
- OpenAI: Preparedness Framework (Beta)2023
Most cited shown; 2 more are mapped here.
Map
Context and risk identification
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models2024
- TrustLLM: Trustworthiness in Large Language Models2024
- Identifying and Mitigating the Security Risks of Generative AI2023
- On the Societal Impact of Open Foundation Models2024
Show 8 moreShow fewer
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory2024
- OWASP AI Security and Privacy Guide2024
- The AI Security Pyramid of Pain2024
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)2024
- An Architectural Risk Analysis of Large Language Models: Applied Machine Learning Security2024
- Generative AI Security: Theories and Practices2024
- OWASP Threat Dragon: AI-Aware Threat Modeling Tool2024
- NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0)2023
Most cited shown; 2 more are mapped here.
Measure
Analyze, assess, and track AI risks
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models2024
- TrustLLM: Trustworthiness in Large Language Models2024
- Identifying and Mitigating the Security Risks of Generative AI2023
Show 8 moreShow fewer
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal2024
- Red-Teaming for Generative AI: Silver Bullet or Security Theater?2024
- Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast2024
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models2024
- SafetyBench: Evaluating the Safety of Large Language Models2024
- A Safe Harbor for AI Evaluation and Red Teaming2024
- AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications2023
- A StrongREJECT for Empty Jailbreaks2024
Most cited shown; 89 more are mapped here.
Manage
Allocate resources to mapped and measured risks
- Constitutional AI: Harmlessness from AI Feedback2022
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned2022
- Identifying and Mitigating the Security Risks of Generative AI2023
- Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly2024
Show 8 moreShow fewer
- Machine Unlearning in Generative AI: A Survey2024
- Lessons From Red Teaming 100 Generative AI Products2025
- OWASP AI Security and Privacy Guide2024
- An Architectural Risk Analysis of Large Language Models: Applied Machine Learning Security2024
- Anthropic's Responsible Scaling Policy2024
- Generative AI Security: Theories and Practices2024
- NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0)2023
- Google: Secure AI Framework (SAIF)2023
ISO/IEC 42001
v2023 · 4 mapped resources