Skip to content
Search
paperSeptember 2026Unreviewed

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch

Abstract

Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures

Categories

Cite

@misc{ranaldi2026worldbench,
  title = {{WorldBench: Culturally Grounded Benchmark for Multilingual Agents}},
  author = {Leonardo Ranaldi and Sherrie Shen and Jushi Kai and Alexandra Birch},
  year = {2026},
  month = sep,
  eprint = {2609.01056},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.01056}
}