June 2026Unreviewed
PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents
Aaditya Pai
Abstract
Prompt injection defenses evaluated on synthetic benchmarks do not generalize to real enterprise documents, which are longer, denser, and interleave legitimate authority language with factual content. We demonstrate this gap with a real-document benchmark of 122 tasks across five professional domains (financial, legal, medical, scientific, DevOps) using actual SEC filings, Federal Register rules, PubMed abstracts, arXiv papers, and GitHub postmortems. Paraphrasing, the strongest defense on synth
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{pai2026parse,
title = {{PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents}},
author = {Aaditya Pai},
year = {2026},
month = jun,
eprint = {2606.17467},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.17467}
}