August 2026Unreviewed
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
Ted Kwartler, Alan Aqrawi, Arian Abbasi
Abstract
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations across many models (Governance Decay; Chen, 2026). We ask a finer question: under a single compaction cycle, how is a safety rule lost, and what does that imply for detection and evaluation? Our central finding is that a presence check is not a safety check: when compaction does
Categories
Cite
@misc{kwartler2026ai,
title = {{AI Guardrail Survival under Single-Cycle Agentic Self-Summarization}},
author = {Ted Kwartler and Alan Aqrawi and Arian Abbasi},
year = {2026},
month = aug,
eprint = {2608.11392},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/d818405f19246444333379662165fb5ee43ebc1e}
}