Skip to content
Search
paperSeptember 2026Unreviewed

Representational alignment yields generalizable safety in language models

Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu

Abstract

Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we sh

Categories

Cite

@misc{li2026representational,
  title = {{Representational alignment yields generalizable safety in language models}},
  author = {Lingyu Li and Yan Teng and Yingchun Wang and Xia Hu},
  year = {2026},
  month = sep,
  eprint = {2609.04022},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.04022}
}