September 2026Unreviewed
Benchmarking LLMs for Threat Level Determination
Han Wang, Murathan Kurfalı, Alfonso Iacovazzi
Abstract
The fast progress of large language models (LLMs) opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a curated dataset derived from publicly available MISP OSINT feeds. Next, we design a tailored prompt to systematically compare eight different LLMs under zero-shot conditions. Finally, we apply supervised fine-tuning on e
Categories
Cite
@misc{wang2026benchmarking,
title = {{Benchmarking LLMs for Threat Level Determination}},
author = {Han Wang and Murathan Kurfalı and Alfonso Iacovazzi},
year = {2026},
month = sep,
eprint = {2609.07582},
archivePrefix = {arXiv},
doi = {10.1109/ICDMW69685.2025.00148},
url = {https://arxiv.org/abs/2609.07582}
}