Skip to content
Search
paperJune 2026Unreviewed

Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks

Zvi Topol

Abstract

Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate (ASR), not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process models from event logs, to red teaming traces. We conduct a controlled experiment pitting 60 HarmBench prompts against two LLMs, GPT-OSS 120B and Llama 3.3 70B, using 10 prompt mutation strategies over

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{topol2026beyond,
  title = {{Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks}},
  author = {Zvi Topol},
  year = {2026},
  month = jun,
  eprint = {2606.07833},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.07833}
}