Updated weekly · 1,386 resources
The research on securing generative AI, in one searchable index.
Papers, standards, tools and reports on prompt injection, jailbreaks, data poisoning and agent security, mapped to OWASP, MITRE ATLAS and NIST.
Recently added
All resources- Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls2026
Dheeraj Mohandas Pai et al.
- Capability-Gated Language Models: Security Composes, Utility Does Not2026
Patrikas Vanagas et al.
- Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective2026
Marco Antonio Corallo et al.
- TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning2026
Muhaimin Bin Munir et al.
- The Safeguard Worked. Is the LLM System Safer?2026
Pingyu Wu et al.
- Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents2026
Chuanchao Zang et al.
Most cited
- ReAct: Synergizing Reasoning and Acting in Language Models2,500cited
Shunyu Yao et al. · ICLR 2023
- Toolformer: Language Models Can Teach Themselves to Use Tools1,400cited
Timo Schick et al. · NeurIPS 2023
- Extracting Training Data from Large Language Models1,200cited
Nicholas Carlini et al. · USENIX Security 2021
- Constitutional AI: Harmlessness from AI Feedback1,100cited
Yuntao Bai et al. · arXiv preprint
- Universal and Transferable Adversarial Attacks on Aligned Language Models890cited
Andy Zou et al. · arXiv preprint
- Voyager: An Open-Ended Embodied Agent with Large Language Models800cited
Guanzhi Wang et al. · NeurIPS 2023
Taxonomy
What the field is working on
Every resource is filed under 46 categories in 8 domains, from attacks and defenses to governance and agentic AI.
Browse all categoriesFramework mappings
Read the research through the standards you use
Know a paper, tool or standard that belongs here?
Suggest it through a GitHub issue. New papers are also collected automatically every week and reviewed before they're merged.