paper/2026arXiv.orgUnreviewed
Hao Li, Ruoyao Wen, Shanghao Shi +2
AI agents that autonomously interact with external tools and environments show great promise across real-world applications. However, the external data which agent consumes also leads to the risk of indirect prompt injection attacks, where malicious instructions embedded in…
paper/2026arXiv.orgUnreviewed
Leo Schwinn, Moritz Ladenburger, Tim Beyer +3
Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to benchmark the robustness of safety…
paper/2026IEEE Transactions on Artificial IntelligenceUnreviewed
Brynn Knowlton, Jovani Campa, Davide Gallo +2
Generative artificial intelligence (AI) systems—particularly large language models (LLMs)—remain vulnerable to jailbreak attacks: adversarial prompts that bypass safeguards and elicit unsafe or restricted outputs. This survey synthesizes jailbreak research from 2023–2025,…
paper/2026Unreviewed
Jialin Song, Xiaodong Liu, Weiwei Yang +4
We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic natural conversational settings, making them easier to bypass safety-aligned LLM than single-turn jailbreaks. Existing…
paper/2026arXiv.orgUnreviewed
Gal Engelberg, Konstantin Koutsyi, Leon Goldberg +5
Identity Security Posture Management (ISPM) is a core challenge for modern enterprises operating across cloud and SaaS environments. Answering basic ISPM visibility questions, such as understanding identity inventory and configuration hygiene, requires interpreting complex…
paper/2026arXiv.orgUnreviewed
I. Steenstra, Paola Pedrelli, Weiyan Shi +2
Large Language Models (LLMs) are increasingly utilized for mental health support; however, current safety benchmarks often fail to detect the complex, longitudinal risks inherent in therapeutic dialogue. We introduce an evaluation framework that pairs AI psychotherapists with…
paper/2026Unreviewed
Michelle Vaccaro, Jaeyoon Song, Abdullah Almaatouq +1
Current frontier AI safety evaluations emphasize static benchmarks, third-party annotations, and red-teaming. In this position paper, we argue that AI safety research should focus on human-centered evaluations that measure harmful capability uplift: the marginal increase in a…
paper/2026arXiv.orgUnreviewed
Zeng Wang, Minghao Shao, Weimin Fu +6
The integration of large language models (LLMs) into electronic design automation (EDA) workflows has introduced powerful capabilities for RTL generation, verification, and design optimization, but also raises critical security concerns. Malicious LLM outputs in this domain pose…
paper/20262026 International Conference on Visual Analytics and Data Visualization (ICVADV)Unreviewed
Akshitha Segireddy
There is a growing use of large language models LLMs in security-related workflows. There is a gap in existing benchmarks for assessing dual-use capabilities of LLMs under safety constraints. We introduce a new, unified benchmark called CybLLM, which encompasses both offensive…
paper/2026Unreviewed
Alexander Nemecek, Osama Zafar, Debargha Ganguly +3
Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English…
paper/2026Unreviewed
P. Wang, Ao-Jie Yuan, Haiyu Zhang +3
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task…
paper/2026InformationUnreviewed
Catalin Anghel, Marian Viorel Craciun, Adina Cocu +4
Large language models are increasingly used as automated graders, yet their reliability under answer-side manipulation and their behavior in multi-model panels remain insufficiently understood. This paper introduces EvalHack, a matrix benchmark in which a fixed committee of four…
paper/2026Unreviewed
Karthik Raghu Iyer, Yazdan Jamshidi, Nicholas Bray +1
We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4$\times$6 Target $\times$ Technique matrix grounded in STRIDE, constructed from a 507-leaf taxonomy -- 401 data-populated and 106 threat-model-derived leaves --…
paper/2026Unreviewed
Zvi Topol
Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that circumvent their safety guardrails. Existing evaluation frameworks typically report binary success/failure metrics, failing to…
paper/2026Unreviewed
Stefan-Claudiu Susan, Andrei Arusoaie, Dorel Lucanu
The irreversible nature of blockchain transactions makes the identification of smart contract vulnerabilities an essential requirement for secure system development. While Large Language Models (LLMs) are increasingly integrated into developer workflows, their reliability as…
paper/2026Unreviewed
Taein Lim, Seongyong Ju, Munhyeok Kim +2
Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit distinct attack patterns. Specifically, each agent exhibits an attack-selection bias,…
paper/2026Unreviewed
Chengjie Wang, Jingzheng Wu, Xiang Ling +2
Large language models (LLMs) are now largely involved in software development workflows, and the code they generate routinely includes third-party library (TPL) imports annotated with specific version identifiers. These version choices can carry security and compatibility risks,…
paper/2026Unreviewed
Haoyu Zhang, Mohammad Zandsalimy, Shanu Sushmita
Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We show that encoding harmful prompts as coherent mathematical problems -- using formalisms such as set theory, formal logic, and…
paper/2026Unreviewed
Maofei Chen, Laifu Wang, Yue Qin +3
How code representation format shapes false positive behaviour in cross-language LLM vulnerability detection remains poorly understood. We systematically vary training intensity and code representation format, comparing raw source text with pruned Abstract Syntax Trees at both…
paper/2026Unreviewed
Sina Mavali, David Pape, Jonathan Evertz +5
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in the environment (e.g., README files, code comments, stack traces) and determine their relevance to…
paper/2026Unreviewed
Chia-Pei, Chen, Kentaroh Toyoda +2
Web-browsing AI agents are increasingly deployed in enterprise settings under strict whitelists of approved domains, yet adversaries can still influence them by embedding hidden instructions in the HTML pages those domains serve. Existing red-teaming resources fall short of this…
paper/2026Unreviewed
Chiyu Zhang, Huiqin Yang, Bendong Jiang +8
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, where an adversary induces an agent to execute dangerous OS-level operations with irreversible…
paper/2026Unreviewed
Xinkai Zhang, Zhipeng Wei, Huanli Gong +4
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can gradually steer the conversation toward an unsafe answer. Recent methods demonstrate this risk, but…
paper/2026Unreviewed
Carsten Maple, Abhishek Kumar, Riya Tapwal
Many jailbreak attack research papers report attack success rates for a limited number of parameter settings, even though there are many combinations of parameter settings that could be used. Further, when new jailbreak papers are released, they often benchmark results against…
paper/2026Unreviewed
Shihao Weng, Yang Feng, Jinrui Zhang +3
The rise of Large Language Model (LLM) agents, augmented with tool use, skills, and external knowledge, has introduced new security risks. Among them, prompt injection attacks, where adversaries embed malicious instructions into the agent workflow, have emerged as the primary…
paper/2026Unreviewed
Scott Thornton
AI-assisted code review is widely used to detect vulnerabilities before production release. Prior work shows that adversarial prompt manipulation can degrade large language model (LLM) performance in code generation. We test whether similar comment-based manipulation misleads…
paper/2026Unreviewed
Strick Sheng, Ziyue Wang, Liyi Zhou
Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations influence tool use, state tracking, and action sequencing, yet their reliability and authority are often uncertain.…
paper/2026Unreviewed
Yanming Mu, Hao Hu, Feiyang Li +7
Retrieval-Augmented Generation (RAG) significantly mitigates the hallucinations and domain knowledge deficiency in large language models by incorporating external knowledge bases. However, the multi-module architecture of RAG introduces complex system-level security…
paper/2026Unreviewed
Chenglin Yang
Modern AI agents execute real-world side effects through tool calls such as file operations, shell commands, HTTP requests, and database queries. A single unsafe action, including accidental deletion, credential exposure, or data exfiltration, can cause irreversible harm.…
paper/2026Unreviewed
Aaron J. Li, Nicolas Sanchez, Hao Huang +8
Large language models (LLMs) are increasingly deployed, yet their outputs can be highly sensitive to routine, non-adversarial variation in how users phrase queries, a gap not well addressed by existing red-teaming efforts. We propose Green Shielding, a user-centric agenda for…
paper/2026Unreviewed
Pedro Yanes Garrido, Diego Fernandez Arias
This paper presents an empirical study on backdoor attacks in large language model agents. We extend a recent attack framework by adding two lightweight benchmarks that measure cross-domain robustness and trigger visibility without changing the model architecture. Our approach…
paper/2026Unreviewed
Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber +2
The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e.,…
paper/2026Unreviewed
He Yang Yuan, Xin Wang, Kundi Yao +3
Logging code plays an important role in software systems by recording key events and behaviors, which are essential for debugging and monitoring. However, insecure logging practices can inadvertently expose sensitive information or enable attacks such as log injection, posing…
paper/2026Unreviewed
Euntae Kim, Soomin Han, Buru Chang
Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to complete, revise, and refine their content. However, this capability poses a serious safety risk: malicious users could jailbreak…
paper/2026Unreviewed
John Pellew, Faizan Raza
How do security scanners perform on real-world code? We present RealVuln, the first open-source benchmark comparing Rule-Based SAST, General-Purpose LLMs, and Security-Specialized scanners on 26 intentionally vulnerable Python repositories (educational and Capture-The-Flag…
paper/2026Unreviewed
Sujan Ghimire, Parsa Mirfasihi, Muhtasim Alam Chowdhury +6
The globalization of integrated circuit (IC) design and manufacturing has increased the exposure of hardware intellectual property (IP) to untrusted stages of the supply chain, raising concerns about reverse engineering, piracy, tampering, and overbuilding. Hardware netlist…
paper/2026Unreviewed
Daniel Zhu, Zihan Wang, Jenny Bao +1
As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this complexity imposes a "jailbreak tax" that degrades the target model's task performance. We show that this tax scales inversely…
paper/2026Unreviewed
Dongcheng Zhang, Yiqing Jiang
Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a distinct and commercially consequential threat category: agents harming their own deployers. Real-world incidents illustrate the…
paper/2026Unreviewed
Harsh Shah
LLM debugging agents that consume cloud logs and execute remediation commands are vulnerable to indirect prompt injection through log content. We present LogJack, a benchmark of 42 payloads across 5 cloud log categories, and evaluate 8 foundation models under 3 prompt conditions…
paper/2026Unreviewed
Qi Li, Jiu Li, Pingtao Wei +8
This report presents a comparative evaluation of DKnownAI Guard in AI agent security scenarios, benchmarked against three competing products: AWS Bedrock Guardrails, Azure Content Safety, and Lakera Guard. Using human annotation as the ground truth, we assess each guardrail's…
paper/2026Unreviewed
Christopher Koch, Joshua Andreas Wellbrock
Agentic AI systems plan, use tools, maintain state, and act across multi-step workflows with external effects, meaning trustworthy deployment can no longer be judged by task completion alone. The current literature remains fragmented across benchmark-centered evaluation,…
paper/2026Unreviewed
Xiaomeng Hu, Yinger Zhang, Fei Huang +7
AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety monitoring to customs import processing), yet existing benchmarks can only evaluate agents in the few domains where public…
paper/2026Unreviewed
Yalun Wu, Haotian Liu, Zhoujun Li +1
As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably reason about complex physics while adhering to safety constraints? We address this through…
paper/2026Unreviewed
Lei Zhao, Abhay Bhaskar, Edgar Dobriban
AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) risk: an agent may execute harmful instructions embedded in untrusted inputs such as email, downloaded files, webpages,…
paper/2026Unreviewed
Yujie Ma, Jialin Rong, Chenxi Yang +4
Large Language Models(LLMs) have been actively integrated into modern software systems as critical components. LLM-in-the-loop vulnerabilities, where vulnerabilities are introduced by LLMs and their dependent downstream components, such as frameworks, introduce new risks.…
paper/2026Unreviewed
Steffen J. Camarato, Yahya Hmaiti, Mandana Ghadamian +1
Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacterized. We present PromptAudit, a controlled evaluation framework that isolates prompt effects by fixing the dataset, decoding, and…
paper/2026Unreviewed
Shi Liu, Xuehai Tang, Xikang Yang +4
The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented autonomous execution capabilities for LLM Agents by integrating external open-domain knowledge and tools. However, this…
paper/2026Unreviewed
Vivek Dahiya, Sunny Nehra, Vipul Dholariya +2
We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118…
paper/2026Unreviewed
Ivan Dobrovolskyi
Organizations that scan documents for sensitive information face a practical problem. Cloud services require data to be sent to external infrastructure, while rule-based tools often miss threats that depend on context. This study presents TorchSight, an open-source local system…
paper/2026Unreviewed
Udari Madhushani Sehwag, Zhengyang Shan, Heming Liu +3
Clarification-seeking behavior is widely regarded as a desirable property of LLM agents, enabling them to resolve ambiguity before acting on underspecified tasks. However, the security implications of this interaction pattern remain unexplored. We investigate whether the…