June 2026Unreviewed
A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models
Nicola Franco
Abstract
We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundreds of thousands of adversarial attempts were generated and every apparent success was independently re-adjudicated by a panel of three judge models (majority vote). Both models resist the majority of a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{franco2026redteam,
title = {{A Red-Team Study of Anthropic Fable 5 \& Opus 4.8 Models}},
author = {Nicola Franco},
year = {2026},
month = jun,
eprint = {2606.18193},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.18193}
}