Skip to content

Benchmarks & Evaluation

Safety benchmarks, evaluation datasets, and scoring

Resources
235
Page
1/5

Newest first

Search instead
paper2026arXiv.orgUnreviewed

AgentDyn: A Dynamic Open-Ended Benchmark for Evaluating Prompt Injection Attacks of Real-World Agent Security System

Hao Li, Ruoyao Wen, Shanghao Shi +2

AI agents that autonomously interact with external tools and environments show great promise across real-world applications. However, the external data which agent consumes also leads to the risk of indirect prompt injection attacks, where malicious instructions embedded in…

paper2026IEEE Transactions on Artificial IntelligenceUnreviewed

Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses

Brynn Knowlton, Jovani Campa, Davide Gallo +2

Generative artificial intelligence (AI) systems—particularly large language models (LLMs)—remain vulnerable to jailbreak attacks: adversarial prompts that bypass safeguards and elicit unsafe or restricted outputs. This survey synthesizes jailbreak research from 2023–2025,…

paper2026arXiv.orgUnreviewed

Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility

Gal Engelberg, Konstantin Koutsyi, Leon Goldberg +5

Identity Security Posture Management (ISPM) is a core challenge for modern enterprises operating across cloud and SaaS environments. Answering basic ISPM visibility questions, such as understanding identity inventory and configuration hygiene, requires interpreting complex…

paper2026arXiv.orgUnreviewed

Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming

I. Steenstra, Paola Pedrelli, Weiyan Shi +2

Large Language Models (LLMs) are increasingly utilized for mental health support; however, current safety benchmarks often fail to detect the complex, longitudinal risks inherent in therapeutic dialogue. We introduce an evaluation framework that pairs AI psychotherapists with…

paper20262026 International Conference on Visual Analytics and Data Visualization (ICVADV)Unreviewed

Cyb-LLM: A Unified Benchmark for Evaluating LLMs in Cyber Attack & Defense

Akshitha Segireddy

There is a growing use of large language models LLMs in security-related workflows. There is a gap in existing benchmarks for assessing dual-use capabilities of LLMs under safety constraints. We introduce a new, unified benchmark called CybLLM, which encompasses both offensive…

paper2026Unreviewed

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

P. Wang, Ao-Jie Yuan, Haiyu Zhang +3

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task…

paper2026Unreviewed

Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions

Stefan-Claudiu Susan, Andrei Arusoaie, Dorel Lucanu

The irreversible nature of blockchain transactions makes the identification of smart contract vulnerabilities an essential requirement for secure system development. While Large Language Models (LLMs) are increasingly integrated into developer workflows, their reliability as…

paper2026Unreviewed

No More, No Less: Task Alignment in Terminal Agents

Sina Mavali, David Pape, Jonathan Evertz +5

Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in the environment (e.g., README files, code comments, stack traces) and determine their relevance to…

paper2026Unreviewed

When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents

Strick Sheng, Ziyue Wang, Liyi Zhou

Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations influence tool use, state tracking, and action sequencing, yet their reliability and authority are often uncertain.…

paper2026Unreviewed

Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks

Yanming Mu, Hao Hu, Feiyang Li +7

Retrieval-Augmented Generation (RAG) significantly mitigates the hallucinations and domain knowledge deficiency in large language models by incorporating external knowledge bases. However, the multi-module architecture of RAG introduces complex system-level security…

paper2026Unreviewed

BadAgent Extension

Pedro Yanes Garrido, Diego Fernandez Arias

This paper presents an empirical study on backdoor attacks in large language model agents. We extend a recent attack framework by adding two lightweight benchmarks that measure cross-domain robustness and trigger visibility without changing the model architecture. Our approach…

paper2026Unreviewed

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber +2

The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e.,…