Skip to content
Search
paperJuly 2026Unreviewed

Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes

Federico Torrielli, Stefano Locci, Amon Rapp, Luigi Di Caro

Abstract

Abstract Large language models are beginning to enter peer review as tools for summarizing manuscripts, drafting evaluations, and reducing reviewer workload. Yet this use creates a security problem specific to evaluative settings: the manuscript being judged can also contain hidden instructions that shape the model’s judgment. We investigate this risk through indirect prompt injection, where hidden text embedded in a manuscript is processed like

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{torrielli2026exploiting,
  title = {{Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes}},
  author = {Federico Torrielli and Stefano Locci and Amon Rapp and Luigi Di Caro},
  year = {2026},
  month = jul,
  doi = {10.1007/s11192-026-05695-x},
  url = {https://doi.org/10.1007/s11192-026-05695-x}
}