2026Unreviewed
ContrastShield: A Contrastive Fine-Tuning Approach for Robust Prompt Injection Detection in LLM Pipelines
Swethaa R
Abstract
Prompt injection attacks pose a critical security threat to large language model (LLM) pipelines, enabling adversaries to hijack model behavior by embedding malicious instructions within user inputs or retrieved tool outputs. Existing defences are primarily rule-based filters and zero-shot classifiers, which are brittle against paraphrased, obfuscated, and indirect injection variants. We present ContrastShield, a binary classifier built on DeBERTa-v3-base and trained with a novel contrastive fin
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{r2026contrastshield,
title = {{ContrastShield: A Contrastive Fine-Tuning Approach for Robust Prompt Injection Detection in LLM Pipelines}},
author = {Swethaa R},
year = {2026},
doi = {10.2139/ssrn.6645520},
url = {https://doi.org/10.2139/ssrn.6645520}
}