A Safe Harbor for AI Evaluation and Red Teaming
Shayne Longpre, Sayash Kapoor, Kevin Klyman, A. Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng-Xin Yong, Suhas Kotha, Yi Zeng, Weiyan Shi, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Ruoxi Jia, Daniel Kang, Sandy Pentland, Arvind Narayanan, Percy Liang, Peter Henderson
International Conference on Machine Learning
Abstract
Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by prominent AI companies to deter model misuse have disincentives on good faith safety evaluations. This causes some researchers to fear that conducting such research or releasing their findings will result in account suspensions or legal reprisal. Although some companies offer researcher access programs, they are an inadequa
Categories
Framework mappings
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@inproceedings{longpre2024safe,
title = {{A Safe Harbor for AI Evaluation and Red Teaming}},
author = {Shayne Longpre and Sayash Kapoor and Kevin Klyman and A. Ramaswami and Rishi Bommasani and Borhane Blili-Hamelin and Yangsibo Huang and Aviya Skowron and Zheng-Xin Yong and Suhas Kotha and Yi Zeng and Weiyan Shi and Xianjun Yang and Reid Southen and Alexander Robey and Patrick Chao and Diyi Yang and Ruoxi Jia and Daniel Kang and Sandy Pentland and Arvind Narayanan and Percy Liang and Peter Henderson},
year = {2024},
month = mar,
booktitle = {International Conference on Machine Learning},
eprint = {2403.04893},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2403.04893},
url = {https://www.semanticscholar.org/paper/21f8977648e25ce8b95020e6f01988af99209c82}
}