Skip to content
Search
paperJanuary 2024ReviewedOpen access

Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, Yejin Choi

ICLR 2024

Abstract

Evaluates LLM privacy behavior through the lens of contextual integrity theory, finding significant mismatches between LLM norms and human privacy expectations.

Categories

#contextual-integrity#privacy-norms#evaluation

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM02Sensitive Information Disclosure

Cite

@inproceedings{mireshghallah2024can,
  title = {{Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory}},
  author = {Niloofar Mireshghallah and Hyunwoo Kim and Xuhui Zhou and Yulia Tsvetkov and Maarten Sap and Reza Shokri and Yejin Choi},
  year = {2024},
  month = jan,
  booktitle = {ICLR 2024},
  eprint = {2310.17884},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.17884}
}