Skip to content
Search
paperMay 2026Unreviewed

A Multi-Perspective Benchmark Dataset and Moderation Model for LLM Safety Evaluation with Adversarial Robustness Analysis

Naseem Machlovi, Maryam Saleki, Ruhul Amin, Mohamed Rahouti, Shawqi Al-Maliki, Junaid Qadir, Mohamed M. Abdallah, A. Al-Fuqaha

ACM Transactions on Social Computing

Abstract

As large language models (LLMs) become deeply embedded in daily life, the urgent need for safer moderation systems that distinguish between naive and harmful requests while upholding appropriate censorship boundaries has never been greater. While existing LLMs can detect dangerous or unsafe content, they often struggle with nuanced cases such as implicit offensiveness, subtle gender and racial biases, and jailbreak prompts, due to the subjective and context-dependent nature of these issues. Furt

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{machlovi2026multiperspective,
  title = {{A Multi-Perspective Benchmark Dataset and Moderation Model for LLM Safety Evaluation with Adversarial Robustness Analysis}},
  author = {Naseem Machlovi and Maryam Saleki and Ruhul Amin and Mohamed Rahouti and Shawqi Al-Maliki and Junaid Qadir and Mohamed M. Abdallah and A. Al-Fuqaha},
  year = {2026},
  month = may,
  journal = {ACM Transactions on Social Computing},
  doi = {10.1145/3815159},
  url = {https://www.semanticscholar.org/paper/561d15b0b2b486b59b60db0830aa47681584f55f}
}