May 2026Unreviewed
OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
Kaixiang Wang, Jiong Lou, Zhaojiacheng Zhou, Jie Li
Abstract
Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce security risks. Existing agentic memory attacks require privileged access or explicit malicious content, making them detectable by advanced safety filters. This leaves a subtler attack surface underexplored: whether adversaries can induce agent to generate experiences that appear locally correct and semantically plausible yet induce harmful general
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{wang2026oep,
title = {{OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences}},
author = {Kaixiang Wang and Jiong Lou and Zhaojiacheng Zhou and Jie Li},
year = {2026},
month = may,
eprint = {2605.18930},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.18930}
}