← Back to search
paper llmsec-2026-00173
Containment Verification: AI Safety Guarantees Independent of Alignment
Royce Moon, Lav R. Varshney
2026-05
Abstract
Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and therefore remain conditional on unverifiable properties of learned behavior. We introduce containment verification, which locates safety guarantees in the agentic framework itself. Under havoc oracle semantics, the AI is modeled as an unconstrained oracle ranging over the entire typed action space, and the verified containment layer must enforce the boundary poli
Categories
Cite This Resource
@article{llmsec202600173,
title = {Containment Verification: AI Safety Guarantees Independent of Alignment},
author = {Royce Moon and Lav R. Varshney},
year = {2026},
url = {https://arxiv.org/abs/2605.09045},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.09045