March 2026Unreviewed
EvalHack: Answer-Side Prompt Injection for Probing LLM Exam-Grading Panel Stability
Catalin Anghel, Marian Viorel Craciun, Adina Cocu, Andreea Alexandra Anghel, Antonio Stefan Balau, Adrian Istrate, Aurelian-Dumitrache Anghele
Information
Abstract
Large language models are increasingly used as automated graders, yet their reliability under answer-side manipulation and their behavior in multi-model panels remain insufficiently understood. This paper introduces EvalHack, a matrix benchmark in which a fixed committee of four LLMs grades university-level machine learning exam answers under a strict integer-only contract (0–10) grounded in instructor-authored rubric artifacts. The dataset comprises 100 students answering 10 short, open-ended i
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@article{anghel2026evalhack,
title = {{EvalHack: Answer-Side Prompt Injection for Probing LLM Exam-Grading Panel Stability}},
author = {Catalin Anghel and Marian Viorel Craciun and Adina Cocu and Andreea Alexandra Anghel and Antonio Stefan Balau and Adrian Istrate and Aurelian-Dumitrache Anghele},
year = {2026},
month = mar,
journal = {Information},
doi = {10.3390/info17030297},
url = {https://doi.org/10.3390/info17030297}
}