Skip to content
Search
paperAugust 2026Unreviewed

NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation

Zhiyuan Xu, Muhammad Firhard Roslan, Joseph Gardiner, Sana Belguith, Lichao Wu

Abstract

Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness. This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same failure outcom

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{xu2026neuronfuzz,
  title = {{NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation}},
  author = {Zhiyuan Xu and Muhammad Firhard Roslan and Joseph Gardiner and Sana Belguith and Lichao Wu},
  year = {2026},
  month = aug,
  eprint = {2608.26222},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.26222}
}