Skip to content
Search
paperMarch 2026Unreviewed

Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift

Michelle Vaccaro, Jaeyoon Song, Abdullah Almaatouq, Michiel A. Bakker

Abstract

Current frontier AI safety evaluations emphasize static benchmarks, third-party annotations, and red-teaming. In this position paper, we argue that AI safety research should focus on human-centered evaluations that measure harmful capability uplift: the marginal increase in a user's ability to cause harm with a frontier model beyond what conventional tools already enable. We frame harmful capability uplift as a core AI safety metric, ground it in prior social science research, and provide concre

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{vaccaro2026evaluating,
  title = {{Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift}},
  author = {Michelle Vaccaro and Jaeyoon Song and Abdullah Almaatouq and Michiel A. Bakker},
  year = {2026},
  month = mar,
  eprint = {2603.26676},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/3a3f4b4c7cd76bff6b5eacacc734dd97d05adc95}
}