Skip to content
Search
paper2026Unreviewed

ContrastShield: A Contrastive Fine-Tuning Approach for Robust Prompt Injection Detection in LLM Pipelines

Swethaa R

Abstract

Prompt injection attacks pose a critical security threat to large language model (LLM) pipelines, enabling adversaries to hijack model behavior by embedding malicious instructions within user inputs or retrieved tool outputs. Existing defences are primarily rule-based filters and zero-shot classifiers, which are brittle against paraphrased, obfuscated, and indirect injection variants. We present ContrastShield, a binary classifier built on DeBERTa-v3-base and trained with a novel contrastive fin

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{r2026contrastshield,
  title = {{ContrastShield: A Contrastive Fine-Tuning Approach for Robust Prompt Injection Detection in LLM Pipelines}},
  author = {Swethaa R},
  year = {2026},
  doi = {10.2139/ssrn.6645520},
  url = {https://doi.org/10.2139/ssrn.6645520}
}