Skip to content
Search
paperMay 2026Unreviewed

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

Shahinul Hoque, Jinghuai Zhang, Jinyuan Sun, Fnu Suya

Abstract

Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit by design: providers hide the model, the tokenizer, and the execution to protect their IP, mitigate jailbreaks, and preserve user privacy, which means an auditor can only inspect proofs the provider supplies. The audit therefore reduces to a consistency check on the provider's o

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hoque2026token,
  title = {{Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage}},
  author = {Shahinul Hoque and Jinghuai Zhang and Jinyuan Sun and Fnu Suya},
  year = {2026},
  month = may,
  eprint = {2605.30040},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.30040}
}