August 2026Unreviewed
WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenance
Jona te Lintelo, Lichao Wu, S. Picek
Abstract
Large Language Model (LLM) watermarks provide a mechanism for text provenance, enabling model owners to identify machine-generated content and attribute it to a specific watermarked model. However, current LLM watermarking approaches predominantly rely on inference-time sampler methods and focus their analysis on dense models. Inference-time methods are only effective when the text is explicitly generated via the model owner's controlled API; they fail in a post-compromise scenario. An adversary
Categories
Cite
@misc{lintelo2026woe,
title = {{WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenance}},
author = {Jona te Lintelo and Lichao Wu and S. Picek},
year = {2026},
month = aug,
eprint = {2608.29151},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/dbfae6a9f6cc8380ca06c4c8bf9f1b077a3bd3ba}
}