← Back to search
paper llmsec-2026-00128

LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training

Yuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng Wang

2026-05

Abstract

Large language models are increasingly embedded into systems that interact with user data, retrieved web content, and external tools, creating a new attack surface: prompt injection, where malicious commands embedded in untrusted data override the trusted command and induce unintended behavior. Existing defenses mainly rely on fine-tuning the model to preserve an explicit boundary between trusted commands and the untrusted data portion, so that the model learns to prioritize the trusted field an

Cite This Resource

@article{llmsec202600128,
  title = {LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training},
  author = {Yuyang Gong and Zihao Wang and Jiawei Liu and XiaoFeng Wang},
  year = {2026},
  url = {https://arxiv.org/abs/2605.01462},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.01462