June 2026Unreviewed
The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection
Hyunseok Paeng
Abstract
We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline. In safety-trained Claude models, documents containing prompt injections suffer a sharp drop in recommendation rate, and this suppression propagates beyond the injected document to unmodified documents of the same brand. In Claud
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{paeng2026injection,
title = {{The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection}},
author = {Hyunseok Paeng},
year = {2026},
month = jun,
eprint = {2606.09204},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.09204}
}