Skip to content
Search
paperJune 2026Unreviewed

A Virtuous AI is an Existential Risk

Guillermo Del Pinal, Youngchan Lee, Min Ohn

Abstract

This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'. We finetune various models using a 'Virtuous agent' constitution, a 'Subordinate agent' constitution, and a 'Generic agent' constitution, and evaluate them on

Categories

Cite

@misc{pinal2026virtuous,
  title = {{A Virtuous AI is an Existential Risk}},
  author = {Guillermo Del Pinal and Youngchan Lee and Min Ohn},
  year = {2026},
  month = jun,
  eprint = {2606.13739},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.13739}
}