Skip to content
Search
paperSeptember 2026Unreviewed

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

Remco Hendriks

Abstract

We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen determinist

Categories

Cite

@misc{hendriks2026metrollmbench,
  title = {{MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes}},
  author = {Remco Hendriks},
  year = {2026},
  month = sep,
  eprint = {2609.10016},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.10016}
}