June 2026Unreviewed
Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense
Minseok Choi, Seungbin Yang, Dongjin Kim, Subin Kim, Jungmin Son, Yunseung Lee, Jaegul Choo, Youngjun Kwak
Abstract
Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks. We propose Membrane, a self-evolving guardrail built on Contrastive Safety Memory (CSM): each cell pairs the conditions for blocking a harmful query with those for permitting a superficially similar ben
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{choi2026membrane,
title = {{Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense}},
author = {Minseok Choi and Seungbin Yang and Dongjin Kim and Subin Kim and Jungmin Son and Yunseung Lee and Jaegul Choo and Youngjun Kwak},
year = {2026},
month = jun,
eprint = {2606.05743},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.05743}
}