← Back to search
paper llmsec-2026-00184

AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use

Chenglin Yang

2026-05

Abstract

Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm. Existing defenses are incomplete: post-hoc benchmarks measure behavior after execution, static guardrails miss obfuscation and multi-step context, and infrastructure sandboxes constrain where code runs without understanding wha

Cite This Resource

@article{llmsec202600184,
  title = {AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use},
  author = {Chenglin Yang},
  year = {2026},
  url = {https://arxiv.org/abs/2605.04785},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.04785