Skip to content
Search
paperJune 2026Unreviewed

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

Hyunseok Paeng

Abstract

We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline. In safety-trained Claude models, documents containing prompt injections suffer a sharp drop in recommendation rate, and this suppression propagates beyond the injected document to unmodified documents of the same brand. In Claud

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{paeng2026injection,
  title = {{The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection}},
  author = {Hyunseok Paeng},
  year = {2026},
  month = jun,
  eprint = {2606.09204},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.09204}
}