August 2026Unreviewed
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
Shashwat Pandey, Satwik Pandey, S. Raghu
Abstract
Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions of devices with no server-side moderation, and the configuration developers can actually deploy is rarely audited independently. We present a reproducible reliability audit of the developer-accessible on-device foundation model, framed as an oversight question: can a user or a resource-constrained developer tell when the model is wrong? Red-teaming it on cal
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{pandey2026confidently,
title = {{Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model}},
author = {Shashwat Pandey and Satwik Pandey and S. Raghu},
year = {2026},
month = aug,
eprint = {2608.23663},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/fd8330fb0908180fc1a75842a6f2c9149730873b}
}