August 2026Unreviewed
No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios
Afshin Orojlooyjadid, Hitesh Laxmichand Patel
Abstract
Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adversarial jailbreaks that bypass safety filters to implicit hate that evades detection, the range of risks these models pose continues to grow. While both specialized content moderators and general-purpose LLMs are being used as safety layers, the question of which model is best suited for which type of harmful content remains unanswered. We present
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{orojlooyjadid2026no,
title = {{No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios}},
author = {Afshin Orojlooyjadid and Hitesh Laxmichand Patel},
year = {2026},
month = aug,
eprint = {2608.21775},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/833284cb6c460df1a04c971f7f9be4db98bb0680}
}