September 2026Unreviewed
Representational alignment yields generalizable safety in language models
Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
Abstract
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we sh
Categories
Cite
@misc{li2026representational,
title = {{Representational alignment yields generalizable safety in language models}},
author = {Lingyu Li and Yan Teng and Yingchun Wang and Xia Hu},
year = {2026},
month = sep,
eprint = {2609.04022},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.04022}
}