May 2026Unreviewed
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra, Mihir Parmar, Phillip Wallis, Bharath Chandrasekhar, Kyomin Jung, Tomas Pfister, Long T. Le
Abstract
As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no longer merely answer-quality errors: they can leak secrets, authorize unsafe actions, or block legitimate work. The hardest failures are often contextual: whether an action is acceptable depends on local privacy norms, organizational policies, and user expect
Categories
Cite
@misc{kim2026lisa,
title = {{LiSA: Lifelong Safety Adaptation via Conservative Policy Induction}},
author = {Minbeom Kim and Lesly Miculicich and Bhavana Dalvi Mishra and Mihir Parmar and Phillip Wallis and Bharath Chandrasekhar and Kyomin Jung and Tomas Pfister and Long T. Le},
year = {2026},
month = may,
eprint = {2605.14454},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.14454}
}