← Back to search
paper llmsec-2026-00128
LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training
Yuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng Wang
2026-05
Abstract
Large language models are increasingly embedded into systems that interact with user data, retrieved web content, and external tools, creating a new attack surface: prompt injection, where malicious commands embedded in untrusted data override the trusted command and induce unintended behavior. Existing defenses mainly rely on fine-tuning the model to preserve an explicit boundary between trusted commands and the untrusted data portion, so that the model learns to prioritize the trusted field an
Categories
Cite This Resource
@article{llmsec202600128,
title = {LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training},
author = {Yuyang Gong and Zihao Wang and Jiawei Liu and XiaoFeng Wang},
year = {2026},
url = {https://arxiv.org/abs/2605.01462},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.01462