Skip to content
Search
paperApril 2026Unreviewed

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

Jiali Wei, Ming Fan, Guoheng Sun, Xicheng Zhang, Haijun Wang, Ting Liu

Abstract

The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studies have demonstrated the feasibility of backdoor attacks against LLMs. However, existing methods suffer from three key shortcomings: explicit trigger patterns that compromise naturalness, unreliable injection of attacker-specified payloads in long-form generation, and incompletely specified threat models that obscure how backdoors are delivered and

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{wei2026stealthy,
  title = {{Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers}},
  author = {Jiali Wei and Ming Fan and Guoheng Sun and Xicheng Zhang and Haijun Wang and Ting Liu},
  year = {2026},
  month = apr,
  eprint = {2604.21700},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.21700}
}