Skip to content
Search
paperApril 2026Unreviewed

Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward

Weiyang Guo, Zesheng Shi, Zeen Zhu, Yuan Zhou, Min Zhang, Jing Li

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical tasks, such as mathematics and programming. However, we identify, for the first time, a latent vulnerability to backdoor attacks within the RLVR framework. This attack can implant a backdoor without modifying the reward verifier by injecting a small amount of poisoning data into the training set. Specifically, we propose a

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{guo2026backdoors,
  title = {{Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward}},
  author = {Weiyang Guo and Zesheng Shi and Zeen Zhu and Yuan Zhou and Min Zhang and Jing Li},
  year = {2026},
  month = apr,
  eprint = {2604.09748},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.09748}
}