Skip to content
Search
paperAugust 2026Unreviewed

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Omatharv Bharat Vaidya, C. Jerzak, Zayne Sprague, Fangcong Yin, N. Hồ

Abstract

Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, a

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{vaidya2026when,
  title = {{When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs}},
  author = {Omatharv Bharat Vaidya and C. Jerzak and Zayne Sprague and Fangcong Yin and N. Hồ},
  year = {2026},
  month = aug,
  eprint = {2608.03506},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/1fd04580be300512828ad334c3a57df3a7b06a63}
}