May 2026Unreviewed
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
Chenglin Yang
Abstract
Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm. Existing defenses are incomplete: post-hoc benchmarks measure behavior after execution, static guardrails miss obfuscation and multi-step context, and infrastructure sandboxes constrain where code runs without understanding wha
Categories
Framework mappings
OWASP Top 10 for Agentic Applications
- ASI02Tool Misuse & Exploitation
MITRE ATLAS
- AML.T0053AI Agent Tool Invocation
Suggested from the entry's categories.
Cite
@misc{yang2026agenttrust,
title = {{AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use}},
author = {Chenglin Yang},
year = {2026},
month = may,
eprint = {2605.04785},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.04785}
}