← Back to search
paper llmsec-2026-00147

Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward

Weiyang Guo, Zesheng Shi, Zeen Zhu, Yuan Zhou, Min Zhang, Jing Li

2026-04

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical tasks, such as mathematics and programming. However, we identify, for the first time, a latent vulnerability to backdoor attacks within the RLVR framework. This attack can implant a backdoor without modifying the reward verifier by injecting a small amount of poisoning data into the training set. Specifically, we propose a

Cite This Resource

@article{llmsec202600147,
  title = {Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward},
  author = {Weiyang Guo and Zesheng Shi and Zeen Zhu and Yuan Zhou and Min Zhang and Jing Li},
  year = {2026},
  url = {https://arxiv.org/abs/2604.09748},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2604.09748