August 2026Unreviewed
When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
Omatharv Bharat Vaidya, C. Jerzak, Zayne Sprague, Fangcong Yin, N. Hồ
Abstract
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{vaidya2026when,
title = {{When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs}},
author = {Omatharv Bharat Vaidya and C. Jerzak and Zayne Sprague and Fangcong Yin and N. Hồ},
year = {2026},
month = aug,
eprint = {2608.03506},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/1fd04580be300512828ad334c3a57df3a7b06a63}
}