August 2026Unreviewed
Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production
Ming Cong, Jingyin Chen, Bin Liu, Qi Chu, Tao Gong, Neng-Hai Yu, Ying-Fei Xiang
Abstract
Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed harmful categories emerge within days, leaving the defense perpetually a step behind. We present SESG (Self-Evolving Safety Guardrails), a multi-agent system running in production. SESG monitors the live traffic behind a deployed guardrail and surfaces two classes of failure: jailbreaks novel in form and harmful categories novel in content. Once
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{cong2026yesterdays,
title = {{Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production}},
author = {Ming Cong and Jingyin Chen and Bin Liu and Qi Chu and Tao Gong and Neng-Hai Yu and Ying-Fei Xiang},
year = {2026},
month = aug,
eprint = {2608.08471},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/ab64ee559216abd0d808ffb581aaba43b7d4825c}
}