Skip to content
Search
paperJune 2026Unreviewed

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

Minseok Choi, Seungbin Yang, Dongjin Kim, Subin Kim, Jungmin Son, Yunseung Lee, Jaegul Choo, Youngjun Kwak

Abstract

Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks. We propose Membrane, a self-evolving guardrail built on Contrastive Safety Memory (CSM): each cell pairs the conditions for blocking a harmful query with those for permitting a superficially similar ben

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{choi2026membrane,
  title = {{Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense}},
  author = {Minseok Choi and Seungbin Yang and Dongjin Kim and Subin Kim and Jungmin Son and Yunseung Lee and Jaegul Choo and Youngjun Kwak},
  year = {2026},
  month = jun,
  eprint = {2606.05743},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.05743}
}