Skip to content
Search
paperAugust 2026Unreviewed

Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model

Shashwat Pandey, Satwik Pandey, S. Raghu

Abstract

Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions of devices with no server-side moderation, and the configuration developers can actually deploy is rarely audited independently. We present a reproducible reliability audit of the developer-accessible on-device foundation model, framed as an oversight question: can a user or a resource-constrained developer tell when the model is wrong? Red-teaming it on cal

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{pandey2026confidently,
  title = {{Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model}},
  author = {Shashwat Pandey and Satwik Pandey and S. Raghu},
  year = {2026},
  month = aug,
  eprint = {2608.23663},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/fd8330fb0908180fc1a75842a6f2c9149730873b}
}