April 2026Unreviewed
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
Weiyang Guo, Zesheng Shi, Zeen Zhu, Yuan Zhou, Min Zhang, Jing Li
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical tasks, such as mathematics and programming. However, we identify, for the first time, a latent vulnerability to backdoor attacks within the RLVR framework. This attack can implant a backdoor without modifying the reward verifier by injecting a small amount of poisoning data into the training set. Specifically, we propose a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{guo2026backdoors,
title = {{Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward}},
author = {Weiyang Guo and Zesheng Shi and Zeen Zhu and Yuan Zhou and Min Zhang and Jing Li},
year = {2026},
month = apr,
eprint = {2604.09748},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.09748}
}