Skip to content
Search
paperApril 2026Unreviewed

PilotBench: A Benchmark for General Aviation Agents with Safety Constraints

Yalun Wu, Haotian Liu, Zhoujun Li, Boyang Wang

Abstract

As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably reason about complex physics while adhering to safety constraints? We address this through PilotBench, a benchmark evaluating LLMs on safety-critical flight trajectory and attitude prediction. Built from 708 real-world general aviation trajectories spanning nine operationally distinct flight phases with synchronized 34-c

Categories

Cite

@misc{wu2026pilotbench,
  title = {{PilotBench: A Benchmark for General Aviation Agents with Safety Constraints}},
  author = {Yalun Wu and Haotian Liu and Zhoujun Li and Boyang Wang},
  year = {2026},
  month = apr,
  eprint = {2604.08987},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.08987}
}