May 2026Unreviewed
Containment Verification: AI Safety Guarantees Independent of Alignment
Royce Moon, Lav R. Varshney
Abstract
Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain conditional on unverifiable properties of learned behavior. We introduce containment verification, which locates safety guarantees in the agentic framework itself. Under havoc oracle semantics, the AI is modeled as an unconstrained oracle ranging over the entire typed action space, and the verified containment layer must enforce the boundary poli
Categories
Cite
@misc{moon2026containment,
title = {{Containment Verification: AI Safety Guarantees Independent of Alignment}},
author = {Royce Moon and Lav R. Varshney},
year = {2026},
month = may,
eprint = {2605.09045},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.09045}
}