January 2024ReviewedOpen access
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, Yejin Choi
ICLR 2024
Abstract
Evaluates LLM privacy behavior through the lens of contextual integrity theory, finding significant mismatches between LLM norms and human privacy expectations.
Categories
#contextual-integrity#privacy-norms#evaluation
Framework mappings
OWASP Top 10 for LLM Applications
- LLM02Sensitive Information Disclosure
NIST AI Risk Management Framework
- GOVERNGovern
- MAPMap
Cite
@inproceedings{mireshghallah2024can,
title = {{Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory}},
author = {Niloofar Mireshghallah and Hyunwoo Kim and Xuhui Zhou and Yulia Tsvetkov and Maarten Sap and Reza Shokri and Yejin Choi},
year = {2024},
month = jan,
booktitle = {ICLR 2024},
eprint = {2310.17884},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2310.17884}
}