Skip to content
Search
paperMay 2026Unreviewed

Containment Verification: AI Safety Guarantees Independent of Alignment

Royce Moon, Lav R. Varshney

Abstract

Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain conditional on unverifiable properties of learned behavior. We introduce containment verification, which locates safety guarantees in the agentic framework itself. Under havoc oracle semantics, the AI is modeled as an unconstrained oracle ranging over the entire typed action space, and the verified containment layer must enforce the boundary poli

Categories

Cite

@misc{moon2026containment,
  title = {{Containment Verification: AI Safety Guarantees Independent of Alignment}},
  author = {Royce Moon and Lav R. Varshney},
  year = {2026},
  month = may,
  eprint = {2605.09045},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.09045}
}