← Back to search
paper llmsec-2026-00147
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
Weiyang Guo, Zesheng Shi, Zeen Zhu, Yuan Zhou, Min Zhang, Jing Li
2026-04
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical tasks, such as mathematics and programming. However, we identify, for the first time, a latent vulnerability to backdoor attacks within the RLVR framework. This attack can implant a backdoor without modifying the reward verifier by injecting a small amount of poisoning data into the training set. Specifically, we propose a
Categories
Cite This Resource
@article{llmsec202600147,
title = {Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward},
author = {Weiyang Guo and Zesheng Shi and Zeen Zhu and Yuan Zhou and Min Zhang and Jing Li},
year = {2026},
url = {https://arxiv.org/abs/2604.09748},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2604.09748