September 2026Unreviewed
Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf
Abstract
Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, we identify 15 systematic failure modes and show that a single pipeline choice can change a model's score by more than 80 percentage points and substantially alter model rankings. A
Categories
Cite
@misc{berriche2026benchmark,
title = {{Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks}},
author = {Aymene Berriche and Cathrine Shalby and Mohannad Alhanahnah and Yazan Boshmaf},
year = {2026},
month = sep,
eprint = {2609.08765},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.08765}
}