Skip to content
Search
paperMay 2026Unreviewed

OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

Kaixiang Wang, Jiong Lou, Zhaojiacheng Zhou, Jie Li

Abstract

Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce security risks. Existing agentic memory attacks require privileged access or explicit malicious content, making them detectable by advanced safety filters. This leaves a subtler attack surface underexplored: whether adversaries can induce agent to generate experiences that appear locally correct and semantically plausible yet induce harmful general

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{wang2026oep,
  title = {{OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences}},
  author = {Kaixiang Wang and Jiong Lou and Zhaojiacheng Zhou and Jie Li},
  year = {2026},
  month = may,
  eprint = {2605.18930},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.18930}
}