August 2026Unreviewed
Learning to Follow In-Context Watermark Instructions via Self-Distillation
Yepeng Liu, Tian-Yi Chen, Xuandong Zhao, D. Song, Yuheng Bu
Abstract
In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model internals. Its reliability hinges on the LLM following the instruction without degrading answer quality, yet how well current LLMs do so has not been measured. We introduce $\mathsf{ICWBench}$, a benchmark of three verifiable ICW instruction families, ea
Categories
Cite
@misc{liu2026learning,
title = {{Learning to Follow In-Context Watermark Instructions via Self-Distillation}},
author = {Yepeng Liu and Tian-Yi Chen and Xuandong Zhao and D. Song and Yuheng Bu},
year = {2026},
month = aug,
eprint = {2608.29030},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/e86439f940d69ec0811550dec8e79a7b1d39a73e}
}