Skip to content
Search
paperJune 2026Unreviewed

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

Seth Dobrin, Łukasz Chmiel

Abstract

AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guardrail libraries. Any control in the agent's address space is reachable by inputs that influence it; this generalizes to any AI system with sufficient reach into its own runtime, a class we term escapable AI systems. We identify four properties that an authorization mecha

Categories

Cite

@misc{dobrin2026unfireable,
  title = {{The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems}},
  author = {Seth Dobrin and Łukasz Chmiel},
  year = {2026},
  month = jun,
  eprint = {2606.26057},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.26057}
}