Skip to content
Search
paperMay 2026Unreviewed

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra, Mihir Parmar, Phillip Wallis, Bharath Chandrasekhar, Kyomin Jung, Tomas Pfister, Long T. Le

Abstract

As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no longer merely answer-quality errors: they can leak secrets, authorize unsafe actions, or block legitimate work. The hardest failures are often contextual: whether an action is acceptable depends on local privacy norms, organizational policies, and user expect

Categories

Cite

@misc{kim2026lisa,
  title = {{LiSA: Lifelong Safety Adaptation via Conservative Policy Induction}},
  author = {Minbeom Kim and Lesly Miculicich and Bhavana Dalvi Mishra and Mihir Parmar and Phillip Wallis and Bharath Chandrasekhar and Kyomin Jung and Tomas Pfister and Long T. Le},
  year = {2026},
  month = may,
  eprint = {2605.14454},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.14454}
}