Skip to content
Search
paperAugust 2026Unreviewed

Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons

Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun

Abstract

Neuron- and path-level interventions offer the finest-grained route to defending large language models (LLMs) against jailbreak attacks, yet existing methods fall short of this promise, i.e., they often compromise model utility significantly. Specifically, one line of work suppresses toxic neurons to erase harmful semantics, but since such semantics are distributed across the network, blocking every pathway forces a large intervention footprint. An alternative line of research focus on identify

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhao2026tripwire,
  title = {{Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons}},
  author = {Wei Zhao and Zhe Li and Peixin Zhang and Jun Sun},
  year = {2026},
  month = aug,
  eprint = {2608.14392},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/49c4e4dfde4fe8e1786ae73918c43796b07e4582}
}