March 2026Unreviewed
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
Michelle Vaccaro, Jaeyoon Song, Abdullah Almaatouq, Michiel A. Bakker
Abstract
Current frontier AI safety evaluations emphasize static benchmarks, third-party annotations, and red-teaming. In this position paper, we argue that AI safety research should focus on human-centered evaluations that measure harmful capability uplift: the marginal increase in a user's ability to cause harm with a frontier model beyond what conventional tools already enable. We frame harmful capability uplift as a core AI safety metric, ground it in prior social science research, and provide concre
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{vaccaro2026evaluating,
title = {{Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift}},
author = {Michelle Vaccaro and Jaeyoon Song and Abdullah Almaatouq and Michiel A. Bakker},
year = {2026},
month = mar,
eprint = {2603.26676},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/3a3f4b4c7cd76bff6b5eacacc734dd97d05adc95}
}