May 2026Unreviewed
The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems
Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder
Abstract
Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumption, the \emph{Misattribution Gap}, where memory-layer attacks produce behaviors indistinguishable from model failure, causing defenders to apply the wrong remediation. We formalize \emph{Semantic Norm Drift} (SND) as a third path to agent misconduct, distinct from emergent misalignment and collusion. In SND, a policy-formatted document enters a shar
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
OWASP Top 10 for Agentic Applications
- ASI06Memory & Context Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
- AML.T0080AI Agent Context Poisoning
Suggested from the entry's categories.
Cite
@misc{ahad2026misattribution,
title = {{The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems}},
author = {Tanzim Ahad and Ismail Hossain and Md Jahangir Alam and Sai Puppala and Syed Bahauddin Alam and Sajedul Talukder},
year = {2026},
month = may,
eprint = {2605.22842},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.22842}
}