Skip to content
Search
paperAugust 2026Unreviewed

COPA: Continual Preference Optimization for Adaptive Prompt Injection Defense

Roshan Sood, Onat Gungor, Tajana Rosing

Abstract

LLMs remain vulnerable to prompt injection attacks, where adversarial instructions embedded in user inputs or external content manipulate model behavior and bypass safeguards. Existing defenses are predominantly static, relying on fixed alignment objectives or attack-specific filtering mechanisms that require redesign as new attack strategies emerge. While recent lifelong alignment methods address shifting user preferences, they do not account for adaptive adversaries that continually evolve to

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{sood2026copa,
  title = {{COPA: Continual Preference Optimization for Adaptive Prompt Injection Defense}},
  author = {Roshan Sood and Onat Gungor and Tajana Rosing},
  year = {2026},
  month = aug,
  eprint = {2608.19982},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.19982}
}