August 2026Unreviewed
Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks
Yuchen Chen, Wei Cheng, Yuan Xiao, Wising Sun, Chunrong Fang, Yang Liu, Zhenyu Chen, Baowen Xu
Abstract
LLM customization platforms allow users to build task-specific models for code intelligence tasks by embedding instructions into system prompts, without modifying the underlying model parameters. While these platforms lower the barrier to developing customized LLMs, they also introduce a new attack surface: instruction backdoor attacks, in which adversaries implant hidden malicious behaviors into customized instructions. However, existing attacks suffer from two key limitations. First, they ofte
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{chen2026breaking,
title = {{Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks}},
author = {Yuchen Chen and Wei Cheng and Yuan Xiao and Wising Sun and Chunrong Fang and Yang Liu and Zhenyu Chen and Baowen Xu},
year = {2026},
month = aug,
eprint = {2608.05659},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.05659}
}