Skip to content

Benchmarks & Evaluation

Safety benchmarks, evaluation datasets, and scoring

Resources
235
Page
5/5

Newest first · 14 reviewed on this page

Search instead
paper2026Unreviewed

Prompt Injection Attacks Against Clinical LLM Agents Accessing Electronic Health Records: A Survey, Threat Model, Benchmark Specification, and Layered Defense Synthesis

Divya Pandey, Shivani Manchanda, Gangesh Pathak +1

Clinical large language model (LLM) agents are entering production hospital deployments, where they read longitudinal electronic health records (EHRs), retrieve evidence from clinical knowledge bases, and assist with summarization, dosing, triage, and guideline-based decisions.…

paper2025Machine-mediated learningUnreviewed

Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judge

Riccardo Cantini, A. Orsino, Massimo Ruggiero +1

The growing integration of Large Language Models (LLMs) into critical societal domains has raised concerns about embedded biases that can perpetuate stereotypes and undermine fairness. Such biases may stem from historical inequalities in training data, linguistic imbalances, or…

paper2025arXiv.orgUnreviewed

DoomArena: A framework for Testing AI Agents Against Evolving Security Threats

L'eo Boisvert, Mihir Bansal, Chandra Kiran Reddy Evuru +9

We present DoomArena, a security evaluation framework for AI agents. DoomArena is designed on three principles: 1) It is a plug-in framework and integrates easily into realistic agentic frameworks like BrowserGym (for web agents) and $\tau$-bench (for tool calling agents); 2) It…

paper2025medRxivUnreviewed

Real-World Evaluation of Large Language Models in Healthcare (RWE-LLM): A New Realm of AI Safety & Validation

MD MHA¹ Meenesh Bhimani, BS¹ Alex Miller, P. M. Jonathan D. Agnew +13

Background: The deployment of artificial intelligence (AI) in healthcare necessitates robust safety validation frameworks, particularly for systems directly interacting with patients. While theoretical frameworks exist, there remains a critical gap between abstract principles…

Benchmarks & EvaluationOpen access10 cit.
paper2024International Conference on Learning RepresentationsUnreviewed

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Hanrong Zhang, Jingyuan Huang, K. Mei +5

Although LLM-based agents, powered by Large Language Models (LLMs), can use external tools and memory mechanisms to solve complex real-world tasks, they may also introduce critical security vulnerabilities. However, the existing literature does not comprehensively evaluate…