September 2026Unreviewed
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
Remco Hendriks
Abstract
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen determinist
Categories
Cite
@misc{hendriks2026metrollmbench,
title = {{MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes}},
author = {Remco Hendriks},
year = {2026},
month = sep,
eprint = {2609.10016},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.10016}
}