September 2026Unreviewed
IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal
Abstract
Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse languages. We introduce IndicSafeEval, a persuasion-based jailbreak evaluation framework for Indian languages. Our benchmark combines ten safety critical content categories with six human-like persuasive strategies across four different Indian languages, such a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{mondal2026indicsafeeval,
title = {{IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks}},
author = {Saikat Mondal and Mamta and Deeksha Varshney and Oana Cocarascu and Asif Ekbal},
year = {2026},
month = sep,
eprint = {2609.03781},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.03781}
}