July 2026Unreviewed
Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes
Federico Torrielli, Stefano Locci, Amon Rapp, Luigi Di Caro
Abstract
Abstract Large language models are beginning to enter peer review as tools for summarizing manuscripts, drafting evaluations, and reducing reviewer workload. Yet this use creates a security problem specific to evaluative settings: the manuscript being judged can also contain hidden instructions that shape the model’s judgment. We investigate this risk through indirect prompt injection, where hidden text embedded in a manuscript is processed like
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{torrielli2026exploiting,
title = {{Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes}},
author = {Federico Torrielli and Stefano Locci and Amon Rapp and Luigi Di Caro},
year = {2026},
month = jul,
doi = {10.1007/s11192-026-05695-x},
url = {https://doi.org/10.1007/s11192-026-05695-x}
}