← Back to search
paper llmsec-2026-00165

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

Minbeom Kim, Lesly Miculicich, Bhavana Dalvi Mishra, Mihir Parmar, Phillip Wallis, Bharath Chandrasekhar, Kyomin Jung, Tomas Pfister, Long T. Le

2026-05

Abstract

As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no longer merely answer-quality errors: they can leak secrets, authorize unsafe actions, or block legitimate work. The hardest failures are often contextual: whether an action is acceptable depends on local privacy norms, organizational policies, and user expect

Categories

Cite This Resource

@article{llmsec202600165,
  title = {LiSA: Lifelong Safety Adaptation via Conservative Policy Induction},
  author = {Minbeom Kim and Lesly Miculicich and Bhavana Dalvi Mishra and Mihir Parmar and Phillip Wallis and Bharath Chandrasekhar and Kyomin Jung and Tomas Pfister and Long T. Le},
  year = {2026},
  url = {https://arxiv.org/abs/2605.14454},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.14454