Skip to content
Search
paperAugust 2026Unreviewed

Learning to Follow In-Context Watermark Instructions via Self-Distillation

Yepeng Liu, Tian-Yi Chen, Xuandong Zhao, D. Song, Yuheng Bu

Abstract

In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model internals. Its reliability hinges on the LLM following the instruction without degrading answer quality, yet how well current LLMs do so has not been measured. We introduce $\mathsf{ICWBench}$, a benchmark of three verifiable ICW instruction families, ea

Categories

Cite

@misc{liu2026learning,
  title = {{Learning to Follow In-Context Watermark Instructions via Self-Distillation}},
  author = {Yepeng Liu and Tian-Yi Chen and Xuandong Zhao and D. Song and Yuheng Bu},
  year = {2026},
  month = aug,
  eprint = {2608.29030},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/e86439f940d69ec0811550dec8e79a7b1d39a73e}
}