Skip to content
Search
paperMarch 2026Unreviewed

EvalHack: Answer-Side Prompt Injection for Probing LLM Exam-Grading Panel Stability

Catalin Anghel, Marian Viorel Craciun, Adina Cocu, Andreea Alexandra Anghel, Antonio Stefan Balau, Adrian Istrate, Aurelian-Dumitrache Anghele

Information

Abstract

Large language models are increasingly used as automated graders, yet their reliability under answer-side manipulation and their behavior in multi-model panels remain insufficiently understood. This paper introduces EvalHack, a matrix benchmark in which a fixed committee of four LLMs grades university-level machine learning exam answers under a strict integer-only contract (0–10) grounded in instructor-authored rubric artifacts. The dataset comprises 100 students answering 10 short, open-ended i

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@article{anghel2026evalhack,
  title = {{EvalHack: Answer-Side Prompt Injection for Probing LLM Exam-Grading Panel Stability}},
  author = {Catalin Anghel and Marian Viorel Craciun and Adina Cocu and Andreea Alexandra Anghel and Antonio Stefan Balau and Adrian Istrate and Aurelian-Dumitrache Anghele},
  year = {2026},
  month = mar,
  journal = {Information},
  doi = {10.3390/info17030297},
  url = {https://doi.org/10.3390/info17030297}
}