April 2026Unreviewed
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
Jiali Wei, Ming Fan, Guoheng Sun, Xicheng Zhang, Haijun Wang, Ting Liu
Abstract
The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studies have demonstrated the feasibility of backdoor attacks against LLMs. However, existing methods suffer from three key shortcomings: explicit trigger patterns that compromise naturalness, unreliable injection of attacker-specified payloads in long-form generation, and incompletely specified threat models that obscure how backdoors are delivered and
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{wei2026stealthy,
title = {{Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers}},
author = {Jiali Wei and Ming Fan and Guoheng Sun and Xicheng Zhang and Haijun Wang and Ting Liu},
year = {2026},
month = apr,
eprint = {2604.21700},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.21700}
}