June 2026Unreviewed
The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents
Manvendra Modgil
Abstract
As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agent have become essential. We study this timing problem using a continuous 18-dimensional affective-dynamics engine (HEART) as a diagnostic probe, evaluating four intervention trigger families - absolute state thresholds, composite state-action patterns, regex reasoning-feature extraction, and zero-shot LLM-as-judge - against human-annotated interv
Categories
Cite
@misc{modgil2026saturation,
title = {{The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents}},
author = {Manvendra Modgil},
year = {2026},
month = jun,
eprint = {2606.04296},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.04296}
}