August 2026Unreviewed
Agent Safety Should Be a Runtime Contract
Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang
Abstract
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases. Agent safety should be a runtime contract enforced by the harness, and the contract has two complementary faces. The preventive face blocks dangerous actions before they happen via sandboxes, permission gates, output filters, and traject
Categories
Cite
@misc{ng2026agent,
title = {{Agent Safety Should Be a Runtime Contract}},
author = {Albus W. Ng and Yi Han and Jusheng Zhang and Wenhao Wang},
year = {2026},
month = aug,
eprint = {2608.11274},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.11274}
}