August 2026Unreviewed
COPA: Continual Preference Optimization for Adaptive Prompt Injection Defense
Roshan Sood, Onat Gungor, Tajana Rosing
Abstract
LLMs remain vulnerable to prompt injection attacks, where adversarial instructions embedded in user inputs or external content manipulate model behavior and bypass safeguards. Existing defenses are predominantly static, relying on fixed alignment objectives or attack-specific filtering mechanisms that require redesign as new attack strategies emerge. While recent lifelong alignment methods address shifting user preferences, they do not account for adaptive adversaries that continually evolve to
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{sood2026copa,
title = {{COPA: Continual Preference Optimization for Adaptive Prompt Injection Defense}},
author = {Roshan Sood and Onat Gungor and Tajana Rosing},
year = {2026},
month = aug,
eprint = {2608.19982},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.19982}
}