Skip to content

Data Poisoning

Training data, fine-tuning, and RAG poisoning attacks

Resources
161
Page
3/4

Newest first

Search instead
paper2026Unreviewed

MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms

Zhen Guo, Shanghao Shi, Shamim Yazdani +2

While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack families. Consequently, whether heterogeneous attacks leave internal representation shifts that…

paper2026Unreviewed

BadEngram: Backdoor Attack on Gated Memory Components in LLMs

Ariel Fogel, Omer Hofman, Eilon Cohen +1

To expand open-weight models' capacity without proportionally increasing computation, recent language models incorporate gated parametric memories that retrieve learned values and inject them into intermediate representations. Despite these efficiency benefits, such modules…

paper2026Unreviewed

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Shuhuai Huang, Jingfeng Zhang, Hong Jia

Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and…

paper2026Unreviewed

Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning

Robin Haselhorst, Lucie Flek, Florian Mai

Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look for. We…

paper2026Security and PrivacyUnreviewed

PERSIST : Threat Modeling Memory‐Persistent AI Agents in Cloud‐to‐Edge Environments

Albert Adusei Brobbey, Narayan P. Bhosale

Agentic artificial intelligence systems increasingly depend on persistent runtime memory, including vector databases, episodic memory stores, long‐term retrieval indices, and cloud‐to‐edge replicas. Existing security frameworks address prompt injection, data poisoning, and…

paper2026Unreviewed

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Xiaodong Wu, Yu Shi, Qi Li +5

Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skill that preserves the payload. We call…

paper2026Unreviewed

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

Jiawei Liu, Jiacheng Guo, Tian Zhang +4

Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks,…

paper2026Unreviewed

Conjunctive Poisoning in AI Supply-Chain Applications

Nokimul Hasan Arif, Qian Lou, Meng Zheng

Large Language and Vision-Language Models are increasingly deployed through inference pipelines that include prompt wrappers (e.g., templates and post-processing scripts) and configuration metadata (e.g., JSON/YAML files) that together shape model outputs. While model weights…

paper2026Kırıkkale Üniversitesi Tıp Fakültesi DergisiUnreviewed

SCENARIO-BASED PREHOSPITAL TRIAGE OF CARBON MONOXIDE POISONING: COMPARING FIVE LARGE LANGUAGE MODELS WITH EMERGENCY MEDICAL SERVICES PERSONNEL IN A PROSPECTIVE STUDY

Vildan Özer, Özlem Bülbül, Efnan Bayrak Erbolukbas +2

Objective: Carbon monoxide (CO) poisoning requires rapid identification and timely decisions regarding the need for hyperbaric oxygen therapy (HBOT) to improve clinical outcomes. This study aimed to compare the decision-making performance of emergency medical services (EMS)…

Data PoisoningOpen access
paper2026Unreviewed

Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

Tao Lin, Gaojie Jin, Zongxi Liu +2

Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers to a finite set of targets before victim training. This assumption substantially underestimates the threat. We show that a single…

paper2026Unreviewed

Backdoor Decontamination Dynamics in LLM Agents

Gabriel Huang, Abhay Puri, L'eo Boisvert +4

Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy…

paper2026Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2Unreviewed

ActivationBackdoor: Backdooring Large Language Models in Collaborative Inference via Intermediate Activations

Zichun Su, Mi Zhang, Xiaohan Zhang +3

Collaborative inference enables cost-effective deployment of large language models by partitioning layers across multiple participants and forwarding intermediate activations between participants in a pipeline, but these transmitted activations also create a new attack surface:…

paper2026Applied SciencesUnreviewed

Design of a Security Framework for Multi-Agent Systems Based on Model Context Protocol in SOC Environments

Rodrigo Tavares de Pina Simões, Xavier Larriva-Novo, Carmen Sánchez-Zas +2

Security Operations Centers (SOCs) rely on Level 1 analysts to triage increasing alert volumes amid alert fatigue and tool fragmentation. LLM-based multi-agent systems using the Model Context Protocol (MCP) are being adopted to automate these tasks, but their autonomy and tool…

paper2026bioRxivUnreviewed

Pervasive Backdoor Vulnerabilities in Genomic Foundation Models

Shiwen Ni, Qianning Wang, Chi Wei +7

Genomic foundation models are increasingly used to interpret and design DNA sequences, yet their susceptibility to training-data manipulation remains poorly understood. Here we systematically evaluate backdoor poisoning across three model families, seven parameter scales ranging…

Data PoisoningOpen access