August 2026Unreviewed
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun
Abstract
Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful semantics, but since such semantics are distributed across the network, blocking every pathway forces a large intervention footprint. An alternative line of research focus on identify
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhao2026tripwire,
title = {{Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons}},
author = {Wei Zhao and Zhe Li and Peixin Zhang and Jun Sun},
year = {2026},
month = aug,
eprint = {2608.14392},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/49c4e4dfde4fe8e1786ae73918c43796b07e4582}
}