May 2026Unreviewed
Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
Zedian Shao, Charles Fleming, Teodora Baluta
Abstract
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that defenses such as outlier detection, clean-data regularization, or online monitoring can neutralize. In this paper, we propose a data poisoning method that teaches an LLM an information hiding scheme reliably and stealthily through semantic associations between shared knowledge such as facts or concepts and attacker-chose
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{shao2026cordyceps,
title = {{Cordyceps: Covert Control Attacks on LLMs via Data Poisoning}},
author = {Zedian Shao and Charles Fleming and Teodora Baluta},
year = {2026},
month = may,
eprint = {2605.26595},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.26595}
}