September 2026Unreviewed
SinoGlyphBench: A Diagnostic Benchmark for Chinese Glyph-Level Obfuscation in Language-Model Moderation
Yifan Wang, Zimu Wang, Suliu Qin, Changyu Zeng, Tong Chen, Siqi Chen, Yijie Lin, Lingyu Jiang, Jionglong Su, Yushan Pan, Haiyang Zhang, Wei Wang, Qiaoyu Tan
Abstract
Glyph-level obfuscation can leave harmful Chinese content readable to humans while degrading automated moderation. We introduce SinoGlyphBench, a diagnostic benchmark that identifies label-critical semantic anchors and creates matched original and glyph-obfuscated inputs in text and image modalities. By perturbing anchors, background context, or both, this design distinguishes corruption of moderation-relevant evidence from general surface variation. Across 176,916 paired evaluations of 12 LLMs
Categories
Cite
@misc{wang2026sinoglyphbench,
title = {{SinoGlyphBench: A Diagnostic Benchmark for Chinese Glyph-Level Obfuscation in Language-Model Moderation}},
author = {Yifan Wang and Zimu Wang and Suliu Qin and Changyu Zeng and Tong Chen and Siqi Chen and Yijie Lin and Lingyu Jiang and Jionglong Su and Yushan Pan and Haiyang Zhang and Wei Wang and Qiaoyu Tan},
year = {2026},
month = sep,
eprint = {2609.05843},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.05843}
}