April 2026Unreviewed
Green Shielding: A User-Centric Approach Towards Trustworthy AI
Aaron J. Li, Nicolas Sanchez, Hao Huang, Ruijiang Dong, Jaskaran Bains, Katrin Jaradeh, Zhen Xiang, Bo Li, Feng Liu, Aaron Kornblith, Bin Yu
Abstract
Large language models (LLMs) are increasingly deployed, yet their outputs can be highly sensitive to routine, non-adversarial variation in how users phrase queries, a gap not well addressed by existing red-teaming efforts. We propose Green Shielding, a user-centric agenda for building evidence-backed deployment guidance by characterizing how benign input variation shifts model behavior. We operationalize this agenda through the CUE criteria: benchmarks with authentic Context, reference standards
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{li2026green,
title = {{Green Shielding: A User-Centric Approach Towards Trustworthy AI}},
author = {Aaron J. Li and Nicolas Sanchez and Hao Huang and Ruijiang Dong and Jaskaran Bains and Katrin Jaradeh and Zhen Xiang and Bo Li and Feng Liu and Aaron Kornblith and Bin Yu},
year = {2026},
month = apr,
eprint = {2604.24700},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.24700}
}