Skip to content
Search
paperAugust 2026Unreviewed

No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios

Afshin Orojlooyjadid, Hitesh Laxmichand Patel

Abstract

Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adversarial jailbreaks that bypass safety filters to implicit hate that evades detection, the range of risks these models pose continues to grow. While both specialized content moderators and general-purpose LLMs are being used as safety layers, the question of which model is best suited for which type of harmful content remains unanswered. We present

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{orojlooyjadid2026no,
  title = {{No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios}},
  author = {Afshin Orojlooyjadid and Hitesh Laxmichand Patel},
  year = {2026},
  month = aug,
  eprint = {2608.21775},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/833284cb6c460df1a04c971f7f9be4db98bb0680}
}