Skip to content
Search
paperSeptember 2026Unreviewed

Benchmarking LLMs for Threat Level Determination

Han Wang, Murathan Kurfalı, Alfonso Iacovazzi

Abstract

The fast progress of large language models (LLMs) opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a curated dataset derived from publicly available MISP OSINT feeds. Next, we design a tailored prompt to systematically compare eight different LLMs under zero-shot conditions. Finally, we apply supervised fine-tuning on e

Categories

Cite

@misc{wang2026benchmarking,
  title = {{Benchmarking LLMs for Threat Level Determination}},
  author = {Han Wang and Murathan Kurfalı and Alfonso Iacovazzi},
  year = {2026},
  month = sep,
  eprint = {2609.07582},
  archivePrefix = {arXiv},
  doi = {10.1109/ICDMW69685.2025.00148},
  url = {https://arxiv.org/abs/2609.07582}
}