Skip to content
Search
paperSeptember 2026Unreviewed

AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis, Yu Feng, Aniruddhan Ramesh, Rico Angell, Shang Hong Sim, Chrysoula Zerva, Emmanouil Koukoumidis

Abstract

LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when a safe solution exists. We introduce AURA-Eval, a framework combining controlled augmentation with granular diagnosis of behavior in tool-use trajectories. Its pipeline identifies safety-critical decision points, generates controlled variations, and constructs counterpa

Categories

Cite

@misc{shang2026auraeval,
  title = {{AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories}},
  author = {Ruoxi Shang and Christina-Maria Androna and Orfeas Menis Mastromichalakis and Yu Feng and Aniruddhan Ramesh and Rico Angell and Shang Hong Sim and Chrysoula Zerva and Emmanouil Koukoumidis},
  year = {2026},
  month = sep,
  eprint = {2609.06783},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06783}
}