Skip to content
Search
paperAugust 2026Unreviewed

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

Ming Cong, Jingyin Chen, Bin Liu, Qi Chu, Tao Gong, Neng-Hai Yu, Ying-Fei Xiang

Abstract

Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed harmful categories emerge within days, leaving the defense perpetually a step behind. We present SESG (Self-Evolving Safety Guardrails), a multi-agent system running in production. SESG monitors the live traffic behind a deployed guardrail and surfaces two classes of failure: jailbreaks novel in form and harmful categories novel in content. Once

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{cong2026yesterdays,
  title = {{Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production}},
  author = {Ming Cong and Jingyin Chen and Bin Liu and Qi Chu and Tao Gong and Neng-Hai Yu and Ying-Fei Xiang},
  year = {2026},
  month = aug,
  eprint = {2608.08471},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/ab64ee559216abd0d808ffb581aaba43b7d4825c}
}