Skip to content
Search
paperAugust 2026UnreviewedOpen access

Data security in large language models: risks, defense, and directions

Kang Chen, Xiu-Ze Zhou, Yuanhui Yu, Y. Lin, Hefeng Chen, Congyu Cai, Li Shen

Journal of King Saud University: Computer and Information Sciences

Abstract

Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine translation, and conversational systems. Despite their transformative potential, these models inherently rely on massive amounts of training data, often collected from diverse and uncurated sources, which exposes them to serious data security risks. Harmful or malicious data can compromise model behavior, leading to toxic outputs or hallucinations, while al

Categories

Cite

@article{chen2026data,
  title = {{Data security in large language models: risks, defense, and directions}},
  author = {Kang Chen and Xiu-Ze Zhou and Yuanhui Yu and Y. Lin and Hefeng Chen and Congyu Cai and Li Shen},
  year = {2026},
  month = aug,
  journal = {Journal of King Saud University: Computer and Information Sciences},
  doi = {10.1007/s44443-026-01200-9},
  url = {https://www.semanticscholar.org/paper/249a53fc8e0e8efa49be0e379ace9460bcd2a0e4}
}