Skip to content

Benchmarks & Evaluation

Safety benchmarks, evaluation datasets, and scoring

Resources
235
Page
3/5

Newest first

Search instead
paper2026Unreviewed

The Poisoned Chalice of LLM Evaluation Report

Jonathan Katzy, Ali Al-Kaswan, Razvan Mihai Popescu +1

Large language models are increasingly used to evaluate and support software engineering tasks, yet the validity of these evaluations is often undermined by uncertainty about whether benchmark instances were seen during pretraining. This can lead to data contamination, which may…

paper2026Unreviewed

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

Shilin Ou, Yifan Xu, Luyao Zhang

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical…

paper2026Unreviewed

NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations

Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain +1

Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that separates untrusted artifact text, trusted…

paper2026Unreviewed

Towards an Automated Test of LLM Security Knowledge

Shufan Chai, Liangliang Sun, Jessica Staddon

Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security…

paper20262026 IEEE 2nd International Conference on Secure IoT, Assured and Trusted Computing (SATC)Unreviewed

Benchmarking the Effectiveness of AI-Driven Red Teaming Across Safety-Aligned Language Models

Joyce Malicha, Kamrul Hasan

Large language models (LLMs) have advanced rapidly, yet even safety-aligned models remain vulnerable to adversarial prompts that bypass safeguards and induce harmful outputs. Conventional red teaming methods, including static testing and gradient-based attacks, are limited by…

paper20262026 International Conference on Connected Intelligence for Industrial Applications (CI2A)Unreviewed

AttaX-Multimodal: End-to-End Evaluation of Multimodal AI Safety with Threatscore and Resiliencescore

Rahul Karne

Currently, there is no single benchmark that can be used to measure the safety and resilience of multimodal AI assistants when subjected to malicious attacks. To fill this gap, we have created AttaX-Multimodal, a comprehensive benchmark of multimodal AI assistant safety that…

paper2026Unreviewed

How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation

Kimberly Milner, Minghao Shao, Nanda Rani +8

Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation…

paper2026Unreviewed

MOLE: Detecting Insider Threats in AI Agents

Aashiq Muhamed, Virginia Smith

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work…