August 2026Unreviewed
Toward a Theory of Value in AI Alignment
Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
Abstract
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without spec
Categories
Cite
@misc{smart2026theory,
title = {{Toward a Theory of Value in AI Alignment}},
author = {Andrew Smart and Shazeda Ahmed and Jackie Kay and Jimmy Tobin and Kris Shrishak and Abeba Birhane},
year = {2026},
month = aug,
eprint = {2608.10327},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.10327}
}