September 2026Unreviewed
MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms
Zhen Guo, Shanghao Shi, Shamim Yazdani, Ning Zhang, Reza Tourani
Abstract
While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack families. Consequently, whether heterogeneous attacks leave internal representation shifts that generalize to unseen threat mechanisms remains unknown. We present MechAudit-40, a systematic evaluation of 40 attack mechanisms across five open-weight model architectures. Threat-specific success criteria, 100,000 matched
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{guo2026mechaudit40,
title = {{MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms}},
author = {Zhen Guo and Shanghao Shi and Shamim Yazdani and Ning Zhang and Reza Tourani},
year = {2026},
month = sep,
eprint = {2609.06612},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.06612}
}