Skip to content
Search
paperAugust 2026Unreviewed

WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenance

Jona te Lintelo, Lichao Wu, S. Picek

Abstract

Large Language Model (LLM) watermarks provide a mechanism for text provenance, enabling model owners to identify machine-generated content and attribute it to a specific watermarked model. However, current LLM watermarking approaches predominantly rely on inference-time sampler methods and focus their analysis on dense models. Inference-time methods are only effective when the text is explicitly generated via the model owner's controlled API; they fail in a post-compromise scenario. An adversary

Categories

Cite

@misc{lintelo2026woe,
  title = {{WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenance}},
  author = {Jona te Lintelo and Lichao Wu and S. Picek},
  year = {2026},
  month = aug,
  eprint = {2608.29151},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/dbfae6a9f6cc8380ca06c4c8bf9f1b077a3bd3ba}
}