Skip to content
Search
paperSeptember 2026Unreviewed

Generative Interpretability via Scalable Neuro-Symbolic Models

Xiaocong Yang

Abstract

As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}

Categories

Cite

@misc{yang2026generative,
  title = {{Generative Interpretability via Scalable Neuro-Symbolic Models}},
  author = {Xiaocong Yang},
  year = {2026},
  month = sep,
  eprint = {2609.13529},
  archivePrefix = {arXiv},
  doi = {10.1145/3806096.3844850},
  url = {https://arxiv.org/abs/2609.13529}
}