← Back to search
paper llmsec-2026-00184
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
Chenglin Yang
2026-05
Abstract
Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm. Existing defenses are incomplete: post-hoc benchmarks measure behavior after execution, static guardrails miss obfuscation and multi-step context, and infrastructure sandboxes constrain where code runs without understanding wha
Cite This Resource
@article{llmsec202600184,
title = {AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use},
author = {Chenglin Yang},
year = {2026},
url = {https://arxiv.org/abs/2605.04785},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.04785