August 2026Unreviewed
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
Tong Zhang, M. Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
Abstract
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer br
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{zhang2026disco,
title = {{DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization}},
author = {Tong Zhang and M. Alfarra and Carlos Hinojosa and Christos Louizos and Bernard Ghanem},
year = {2026},
month = aug,
eprint = {2608.17067},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/74f0b8c1be197c2e6a6d54d29c20352d640cb979}
}