Skip to content
Search
paperSeptember 2026Unreviewed

MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms

Zhen Guo, Shanghao Shi, Shamim Yazdani, Ning Zhang, Reza Tourani

Abstract

While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack families. Consequently, whether heterogeneous attacks leave internal representation shifts that generalize to unseen threat mechanisms remains unknown. We present MechAudit-40, a systematic evaluation of 40 attack mechanisms across five open-weight model architectures. Threat-specific success criteria, 100,000 matched

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{guo2026mechaudit40,
  title = {{MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms}},
  author = {Zhen Guo and Shanghao Shi and Shamim Yazdani and Ning Zhang and Reza Tourani},
  year = {2026},
  month = sep,
  eprint = {2609.06612},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06612}
}