September 2026Unreviewed
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
Abstract
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors
Categories
Cite
@misc{abootorabi2026steering,
title = {{Steering Geometry: Validating Human Value Geometry in LLM Steering Space}},
author = {Mohammad Mahdi Abootorabi and Armin Saghafian and Ali Bazshoushtari and Hamid Rezaei and EunJeong Hwang and Vered Shwartz and Parvin Mousavi and Purang Abolmaesumi},
year = {2026},
month = sep,
eprint = {2609.06289},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.06289}
}