Skip to content

Adversarial Examples

Evasion attacks and adversarial perturbations

Resources
78
Page
1/2

Newest first

Search instead
paper2026International Journal of Innovative Research and Creative TechnologyUnreviewed

Adversarial Machine Learning Threats To Medical Device AI Controllers

Venkata Sai Abhinav Piratla -

The integration of artificial intelligence into life-critical medical device controllers—including closed-loop insulin delivery systems and cardiac monitoring devices—introduces adversarial machine learning (AML) attack surfaces that conventional cybersecurity frameworks do not…

paper2026Unreviewed

Laundering AI Authority with Adversarial Examples

Jie Zhang, Pura Peetathawatchai, Florian Tramèr +1

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial…

paper2026Unreviewed

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training

Yuyang Gong, Zihao Wang, Jiawei Liu +1

Large language models are increasingly embedded into systems that interact with user data, retrieved web content, and external tools, creating a new attack surface: prompt injection, where malicious commands embedded in untrusted data override the trusted command and induce…

paper2026Unreviewed

Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks

Yanming Mu, Hao Hu, Feiyang Li +7

Retrieval-Augmented Generation (RAG) significantly mitigates the hallucinations and domain knowledge deficiency in large language models by incorporating external knowledge bases. However, the multi-module architecture of RAG introduces complex system-level security…

paper2026Unreviewed

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber +2

The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This paper evaluates the risk of jailbreaking LLMs, i.e.,…

paper2026Unreviewed

Attention Is Where You Attack

Aviral Srivastava, Sourav Panda

Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety behavior remain poorly understood. We introduce the Attention Redistribution Attack (ARA), a white-box adversarial attack that…

paper2026Unreviewed

Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures

Faisal Haque Bappy, Tahrim Hossain, Tarannum Shaila Zaman +3

Multi-agent LLM pipelines orchestrate multiple specialized language model agents into structured workflows where intermediate outputs are passed across agents to solve complex tasks. This design introduces a security gap absent in single-agent settings: once an agent accepts…

paper20262026 International Conference on Connected Intelligence for Industrial Applications (CI2A)Unreviewed

AttaX-Multimodal: End-to-End Evaluation of Multimodal AI Safety with Threatscore and Resiliencescore

Rahul Karne

Currently, there is no single benchmark that can be used to measure the safety and resilience of multimodal AI assistants when subjected to malicious attacks. To fill this gap, we have created AttaX-Multimodal, a comprehensive benchmark of multimodal AI assistant safety that…

paper2026ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)Unreviewed

EGP-Defense: Enhancing Adversarial Robustness of LVLMs via Training-Free Edge-Guided Prompting

Bo-Yu Wang, Zi-Wen He, Xin-Jue Hu +3

Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal comprehension capabilities, achieving state-of-the-art performance across various vision-language tasks. However, their performance drops significantly when facing adversarial attacks on the visual…

paper2026IEEE Transactions on Dependable and Secure ComputingUnreviewed

Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth

Hai-Qi Zhang, Zi-Qiang Li, Hao Tang +1

Vision-Language Pre-training (VLP) models are highly capable in multimodal tasks but are critically vulnerable to adversarial attacks. Existing methods for creating transferable adversarial examples typically operate by modifying semantics within isolated image-text pairs. This…

paper2026Unreviewed

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

Jiawei Liu, Jiacheng Guo, Tian Zhang +4

Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks,…