Skip to content
Search
paperMay 2026Unreviewed

AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use

Chenglin Yang

Abstract

Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm. Existing defenses are incomplete: post-hoc benchmarks measure behavior after execution, static guardrails miss obfuscation and multi-step context, and infrastructure sandboxes constrain where code runs without understanding wha

Categories

Framework mappings

OWASP Top 10 for Agentic Applications
  • ASI02Tool Misuse & Exploitation
MITRE ATLAS
  • AML.T0053AI Agent Tool Invocation

Suggested from the entry's categories.

Cite

@misc{yang2026agenttrust,
  title = {{AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use}},
  author = {Chenglin Yang},
  year = {2026},
  month = may,
  eprint = {2605.04785},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.04785}
}