Skip to content
Search
paperMay 2026Unreviewed

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder

Abstract

Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumption, the \emph{Misattribution Gap}, where memory-layer attacks produce behaviors indistinguishable from model failure, causing defenders to apply the wrong remediation. We formalize \emph{Semantic Norm Drift} (SND) as a third path to agent misconduct, distinct from emergent misalignment and collusion. In SND, a policy-formatted document enters a shar

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
OWASP Top 10 for Agentic Applications
  • ASI06Memory & Context Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data
  • AML.T0080AI Agent Context Poisoning

Suggested from the entry's categories.

Cite

@misc{ahad2026misattribution,
  title = {{The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems}},
  author = {Tanzim Ahad and Ismail Hossain and Md Jahangir Alam and Sai Puppala and Syed Bahauddin Alam and Sajedul Talukder},
  year = {2026},
  month = may,
  eprint = {2605.22842},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.22842}
}