July 2026Unreviewed
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety
Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra
Abstract
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{ifebi2026minionese,
title = {{Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety}},
author = {Chigozirim Ifebi and Brent Kong and Ayushi Mehrotra},
year = {2026},
month = jul,
eprint = {2607.10112},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.10112}
}