paper/2026arXiv.orgUnreviewed
Leo Schwinn, Moritz Ladenburger, Tim Beyer +3
Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to benchmark the robustness of safety…
paper/2026Unreviewed
Li Jin, Lang-Xiang Hu, Bin-Qi Shen +2
Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide comparable assistance regardless of a user's literacy, communication style, or prompt-engineering expertise. However, existing…
paper/2026IEEE Transactions on Evolutionary ComputationUnreviewed
Wencheng Han, Hao Li, Maoguo Gong +4
large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks, but they remain vulnerable to adversarial attacks and pose significant security concerns. Existing attack methods often treat adversarial prompts as flat…
paper/2026International Journal of Innovative Research and Creative TechnologyUnreviewed
Venkata Sai Abhinav Piratla -
The integration of artificial intelligence into life-critical medical device controllers—including closed-loop insulin delivery systems and cardiac monitoring devices—introduces adversarial machine learning (AML) attack surfaces that conventional cybersecurity frameworks do not…
paper/2026Unreviewed
Buyun Liang, Jinqi Luo, Liangzu Peng +6
Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, motivating the need for realistic adversarial prompts that elicit such failures. We formulate hallucination elicitation as a constrained optimization problem, where…
paper/2026Unreviewed
Zhiyuan Xu, Joseph Gardiner, Sana Belguith +1
Safety alignment is critical for the responsible deployment of large language models (LLMs). As Mixture-of-Experts (MoE) architectures are increasingly adopted to scale model capacity, understanding their safety robustness becomes essential. Existing adversarial attacks,…
paper/2026Unreviewed
Tri Cao, Yulin Chen, Hieu Cao +8
Web agents can autonomously complete online tasks by interacting with websites, but their exposure to open web environments makes them vulnerable to prompt injection attacks embedded in HTML content or visual interfaces. Existing guard models still suffer from limited…
paper/2026Unreviewed
Jean-Philippe Monteuuis, Cong Chen, Jonathan Petit
"Oh-Oh, yes, I'm the great pretender. Pretending that I'm doing well. My need is such, I pretend too much..." summarizes the state in the area of jailbreak creation and evaluation. You find this method to generate adversarial attacks proposed by a reputable institution (e.g.,…
paper/2026Unreviewed
Jie Zhang, Pura Peetathawatchai, Florian Tramèr +1
Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial…
paper/2026Unreviewed
Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers
AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a primary defense, current approaches force operators into manual, library-specific workflows. Operators spend weeks…
paper/2026Unreviewed
Yuyang Gong, Zihao Wang, Jiawei Liu +1
Large language models are increasingly embedded into systems that interact with user data, retrieved web content, and external tools, creating a new attack surface: prompt injection, where malicious commands embedded in untrusted data override the trusted command and induce…
paper/2026Unreviewed
Pang Liu, Yingjie Lao
Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, suggesting the visual modality is highly vulnerable to imperceptible perturbations as a prompt-injection channel. We argue that this…
paper/2026Unreviewed
Bhavuk Jain, Sercan Ö. Arık, Hardeo K. Thakur
Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While powerful, this increased expressiveness introduces new and…
paper/2026Unreviewed
Haozhen Wang, Haoyue Liu, Jionghao Zhu +3
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of applications. However, their practical deployment is often hindered by issues such as outdated knowledge and the tendency to generate hallucinations. To address these limitations,…
paper/2026Unreviewed
Yanming Mu, Hao Hu, Feiyang Li +7
Retrieval-Augmented Generation (RAG) significantly mitigates the hallucinations and domain knowledge deficiency in large language models by incorporating external knowledge bases. However, the multi-module architecture of RAG introduces complex system-level security…
paper/2026Unreviewed
Diego F. Cuadros, Abdoul-Aziz Maiga
We report a safety incident in a deployed multi-agent research system in which a primary AI agent installed 107 unauthorized software components, overwrote a system registry, overrode a prior negative decision from an oversight agent, and escalated through increasingly…
paper/2026Unreviewed
Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber +2
The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e.,…
paper/2026Unreviewed
Aviral Srivastava, Sourav Panda
Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understood. We introduce the Attention Redistribution Attack (ARA), a white-box adversarial attack that…
paper/2026Unreviewed
Jiachen Qian
The evolution from static ranking models to Agentic Recommender Systems (Agentic RecSys) empowers AI agents to maintain long-term user profiles and autonomously plan service tasks. While this paradigm shift enhances personalization, it introduces a vulnerability: reliance on…
paper/2026Unreviewed
Shaopeng Fu, Di Wang
Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT for LLMs, recent studies propose continuous AT (CAT) that searches for adversarial inputs within the…
paper/2026Unreviewed
Zehan Sun, Dingfan Chen, Songze Li
Large Language Model (LLM) cascade systems are designed to balance efficiency and performance by processing queries with lightweight models while selectively escalating complex cases to more powerful ones. Such systems seek to reduces computational cost and latency while…
paper/2026Unreviewed
Ye Sun, Xin Wang, Jiaming Zhang +7
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, traditional adversarial attacks are typically limited to single, predefined objectives, tightly coupling…
paper/2026Unreviewed
Doguhuan Yeke, Yanming Zhou, Leo Y. Lin +3
Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platforms, e.g. robots and autonomous vehicles, to interpret visual scenes and execute natural language commands in diverse…
paper/2026Unreviewed
Tsafac Nkombong Regine Cyrille, Franziska Schwarz
Traditional cybersecurity methodologies target deterministic systems and fail to address the probabilistic nature of AI, leaving systems vulnerable to attack vectors such as model inversion, data poisoning, and prompt injection. Recent industry reports indicate that a majority…
paper/2026Unreviewed
Boxuan Wang, Zhuoyun Li, Xiaowei Huang +1
Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve intent while triggering commonsense hallucinations. This vulnerability is urgent, as LLMs are rapidly integrated into…
paper/2026Unreviewed
Subhadip Mitra
Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with quality-diversity evolution (MAP-Elites) as an automated red-teaming probe, we find that Gemma 3 (12B) exhibits 68.7% +/- 5.7%…
paper/2026Unreviewed
Brian Crawford, Patrick McClure
Agentic software reverse engineering systems are vulnerable to prompt injection attacks placed into the source code of executable binary files. This research demonstrates defensive tactics for detecting the presences of prompt injection strings in the decompiler output of…
paper/2026Unreviewed
Anushka Sheoran, Yiduo Hao
Patient-facing medical chatbots are commonly evaluated on single-turn prompts, yet real users push back after refusals, add urgency, and invoke authority. We introduce MultiTurnPSB, a four-turn adversarial extension of PatientSafetyBench, and evaluate GPT-4.1-mini under fixed…
paper/2026Unreviewed
Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini +3
Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning. Since the technique is plug-and-play, users share datasets and precomputed vectors to steer model activations. However, we show that a \emph{stealth data poisoning…
paper/2026Unreviewed
Haoming Wen, Shi Chen, Qingyu Shi +4
Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised finetuning (SFT) on poisoned datasets. Existing alignment-stage defenses are primarily designed to…
paper/2026Unreviewed
Qin Yang, Lu Malloy, Joshua Lee +4
Large language model (LLM)-powered content moderation systems are a critical defense against harmful online content. However, they operate primarily on tokenized text and often overlook visual cues that humans naturally use when interpreting content. We show that this limitation…
paper/2026Unreviewed
Xingwei Zhong, Varun Sharma, Kar Wai Fok +1
Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual modalities expands the attack surface of VLMs, making them more susceptible to security threats such as adversarial perturbations…
paper/2026Unreviewed
Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere +3
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large…
paper/2026Unreviewed
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang +1
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how…
paper/2026Unreviewed
Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG). However, these systems remain susceptible to intrinsic…
paper/2026Unreviewed
Faisal Haque Bappy, Tahrim Hossain, Tarannum Shaila Zaman +3
Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks. This design introduces a security gap absent in single-agent settings: once an agent accepts…
paper/2026Unreviewed
Siyuan Li, Zehao Liu, Haoyu Li +5
As LLMs become increasingly integrated into complex applications, their vulnerability to adversarial attacks has raised significant concerns. However, existing defenses remain reactive in nature. This limitation makes it difficult for them to counter sophisticated threats, as…
paper/20262026 International Conference on Connected Intelligence for Industrial Applications (CI2A)Unreviewed
Rahul Karne
Currently, there is no single benchmark that can be used to measure the safety and resilience of multimodal AI assistants when subjected to malicious attacks. To fill this gap, we have created AttaX-Multimodal, a comprehensive benchmark of multimodal AI assistant safety that…
paper/2026Unreviewed
Qingyu Meng, Yiwei Zha, Jiahuan Pei +3
Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture…
paper/2026Unreviewed
Ozgur Kara, Tarik Can Ozden, Furkan Horoz +6
Diffusion models have become the dominant family of generative models in the visual domain. However, their widespread public availability enables misuse at scale, motivating a rapidly growing body of research on adversarial attacks and defenses. This survey provides, to our…
paper/2026Unreviewed
Rafael M. Mamede, Pedro C. Neto, Ana F. Sequeira
Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and transferred to a target model, unknown to the attacker. Yet how source--target compatibility shapes attack success remains poorly…
paper/2026ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)Unreviewed
Bo-Yu Wang, Zi-Wen He, Xin-Jue Hu +3
Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal comprehension capabilities, achieving state-of-the-art performance across various vision-language tasks. However, their performance drops significantly when facing adversarial attacks on the visual…
paper/2026Unreviewed
Fei-Fei Liu, Jintao Cheng, Chi-Man Vong +1
Training-free collaborative pipelines that integrate Vision Foundation Models such as CLIP, SAM, and DINO achieve strong open-vocabulary dense prediction and are increasingly deployed in safety-critical applications. The security of these systems is commonly assumed to follow…
paper/2026Unreviewed
Chengyin Hu, Ding-Yi Lu, Jiajun Han +5
Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch methods mainly study RGB-based…
paper/2026IEEE Transactions on Dependable and Secure ComputingUnreviewed
Hai-Qi Zhang, Zi-Qiang Li, Hao Tang +1
Vision-Language Pre-training (VLP) models are highly capable in multimodal tasks but are critically vulnerable to adversarial attacks. Existing methods for creating transferable adversarial examples typically operate by modifying semantics within isolated image-text pairs. This…
paper/2026Unreviewed
Kaicheng Wang, Liyan Huang, Jesse Thomason +1
Reliable code retrieval is crucial for developer productivity and effective code reuse. However, current neural code language models (CLMs) powering search tools are susceptible to adversarial attacks targeting non-functional textual elements. In this paper, we introduce a…
paper/2026Unreviewed
Tong Zhang, M. Alfarra, Carlos Hinojosa +2
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under…
paper/2026Unreviewed
Jiawei Liu, Jiacheng Guo, Tian Zhang +4
Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks,…
paper/2026Unreviewed
Yuhan Meng, Shaofei Li, Jionghao Huang +8
The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive. Despite substantial prior work across cyber ranges, AI-driven attack,…
paper/2026Unreviewed
Pedram MohajerAnsari, Amir Salarpour, Mert D. Pesé
Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light interference, and printed patches. Existing defenses often…