Skip to content
Search
paperMarch 2026Unreviewed

MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

Zhong-Qiu Wang, Yueqian Lin, Jingyang Zhang, Hai Li, Yiran Chen

Abstract

Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, run-centric platform that integrates automatic cross-modal payload generation, three multi-turn attack algorithms (Crescendo, PAIR, Violent Durian), provider-agnostic model routing, and an LLM judge

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{wang2026muse,
  title = {{MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models}},
  author = {Zhong-Qiu Wang and Yueqian Lin and Jingyang Zhang and Hai Li and Yiran Chen},
  year = {2026},
  month = mar,
  eprint = {2603.02482},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/8d3675bff32bbde8709b044d721dc88a7ca55c8b}
}