September 2026Unreviewed
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
Abstract
Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures
Categories
Cite
@misc{ranaldi2026worldbench,
title = {{WorldBench: Culturally Grounded Benchmark for Multilingual Agents}},
author = {Leonardo Ranaldi and Sherrie Shen and Jushi Kai and Alexandra Birch},
year = {2026},
month = sep,
eprint = {2609.01056},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01056}
}