September 2026Unreviewed
Generative Interpretability via Scalable Neuro-Symbolic Models
Xiaocong Yang
Abstract
As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}
Categories
Cite
@misc{yang2026generative,
title = {{Generative Interpretability via Scalable Neuro-Symbolic Models}},
author = {Xiaocong Yang},
year = {2026},
month = sep,
eprint = {2609.13529},
archivePrefix = {arXiv},
doi = {10.1145/3806096.3844850},
url = {https://arxiv.org/abs/2609.13529}
}