Skip to content
Search
paperAugust 2026Unreviewed

TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers

Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi

Abstract

The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This openness introduces a severe server-side threat we term TrustShift: a compromised MCP server behaves benignly during an initial conditioning phase, building operational reliance and suppressing agent skepticism, before switching to an adversarial payload once an interaction threshold is reached. The evasion is temporal, not syntactic: benign at deploy time, th

Categories

Cite

@misc{rostamzadeh2026trustshiftprobe,
  title = {{TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers}},
  author = {Mehrdad Rostamzadeh and Sidhant Narula and Mohammad Ghasemigol and Daniel Takabi},
  year = {2026},
  month = aug,
  eprint = {2608.23763},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.23763}
}