April 2024ReviewedOpen access
Anthropic: Many-shot Jailbreaking
Anthropic
Anthropic Research Blog
Abstract
Reveals many-shot jailbreaking, a technique exploiting long context windows by including many examples of harmful Q&A pairs to override safety training.
Categories
#many-shot#long-context#in-context-learning
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Cite
@techreport{anthropic2024anthropic,
title = {{Anthropic: Many-shot Jailbreaking}},
author = {{Anthropic}},
year = {2024},
month = apr,
institution = {Anthropic Research Blog},
url = {https://www.anthropic.com/research/many-shot-jailbreaking}
}