Skip to content
Search
paperJune 2026Unreviewed

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

Malia Barker, Bishal Lakha, Edoardo Serra, Francesco Gullo

Abstract

Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate computation to code. Yet models are still often used in settings where they must reason directly from natural language, and trustworthy models should solve small-number arithmetic word problems without external tools. Prior work shows that LLMs are sensitive to numerical variation: a model may solve an original problem but fail on structurally simi

Categories

Cite

@misc{barker2026testing,
  title = {{Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks}},
  author = {Malia Barker and Bishal Lakha and Edoardo Serra and Francesco Gullo},
  year = {2026},
  month = jun,
  eprint = {2606.03606},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.03606}
}