Skip to content
Search
paperAugust 2026Unreviewed

Agent Safety Should Be a Runtime Contract

Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang

Abstract

The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases. Agent safety should be a runtime contract enforced by the harness, and the contract has two complementary faces. The preventive face blocks dangerous actions before they happen via sandboxes, permission gates, output filters, and traject

Categories

Cite

@misc{ng2026agent,
  title = {{Agent Safety Should Be a Runtime Contract}},
  author = {Albus W. Ng and Yi Han and Jusheng Zhang and Wenhao Wang},
  year = {2026},
  month = aug,
  eprint = {2608.11274},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.11274}
}