← Back to search
paper llmsec-2026-00179

Automated alignment is harder than you think

Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau, Geoffrey Irving

2026-05

Abstract

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tas

Categories

Cite This Resource

@article{llmsec202600179,
  title = {Automated alignment is harder than you think},
  author = {Aleksandr Bowkis and Marie Davidsen Buhl and Jacob Pfau and Geoffrey Irving},
  year = {2026},
  url = {https://arxiv.org/abs/2605.06390},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.06390