Skip to content

Benchmarks & Evaluation

Safety benchmarks, evaluation datasets, and scoring

Resources
235
Page
4/5

Newest first

Search instead
paper2026Unreviewed

Benchmarking LLMs for Threat Level Determination

Han Wang, Murathan Kurfalı, Alfonso Iacovazzi

The fast progress of large language models (LLMs) opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a…

paper2026Unreviewed

VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities

Jiahao Shi, Edward Tsien, Yifeng Di +10

The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a…

paper2026Unreviewed

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif +2

Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that accuracy reflects learning or target leakage, whether tabular foundation models change the answer,…

paper2026International Journal For Multidisciplinary ResearchUnreviewed

Adversarial Robustness of Foundation Models for Intelligent Mechanical Systems: Threat Models, Benchmarks, and Defense Stacks

Vishwanath

Foundation models increasingly operate across modalities (vision, language, audio, and vision–language) and are deployed in decision-critical pipelines with tool use and retrieval. This expands the adversarial surface: small perturbations to images or audio can flip predictions,…

paper2026Journal of Cybersecurity, Digital Forensics and JurisprudenceUnreviewed

Toward Dynamic and Risk-Aware Evaluation of Cybersecurity LLMs: A Survey and the RIRAG Framework

Ravi Prasad, Feroz Ahmed, Shohel Rana +3

The rapid adoption of large language models (LLMs) in cybersecurity has created a growing need for evaluation methods that reflect operational risk rather than isolated language capability. Existing cybersecurity benchmarks assess useful dimensions such as factual knowledge,…

paper2026Yalvaç akademi dergisiUnreviewed

Leakage-Aware Cross-Dataset Evaluation of Prompt Injection Detection Using Classical Machine Learning and Transformer Models

Oğuzhan Kilim

The widespread adoption of systems based on Large Language Models has made the reliable detection of prompt injection attacks a critical requirement. However, high performance achieved on training and test splits generated from the same data source does not guarantee that models…

paper2026ComputersUnreviewed

Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication

Adil Khan, Khaled AlKhanbashi, Azza Mohamed

Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4,…

paper2026BMC Medical Informatics and Decision MakingUnreviewed

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

Sanjay Basu, Sadiq Y. Patel, Parth Sheth +4

Studies of artificial intelligence tools used in patient triage have largely involved academic medical center cohorts, scripted patient-actor scenarios, or knowledge benchmarks. Populations that may rely on such tools due to constrained access to in-person care, including…

paper2026Unreviewed

Language-Specific Gaps in AI Safety Training Datasets

Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru

Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage claims frequently do not survive inspection at the…