June 2026Unreviewed
The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems
Seth Dobrin, Łukasz Chmiel
Abstract
AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guardrail libraries. Any control in the agent's address space is reachable by inputs that influence it; this generalizes to any AI system with sufficient reach into its own runtime, a class we term escapable AI systems. We identify four properties that an authorization mecha
Categories
Cite
@misc{dobrin2026unfireable,
title = {{The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems}},
author = {Seth Dobrin and Łukasz Chmiel},
year = {2026},
month = jun,
eprint = {2606.26057},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.26057}
}