August 2026Unreviewed
TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi
Abstract
The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model agents to external tool backends. This openness introduces a severe server-side threat we term TrustShift: a compromised MCP server behaves benignly during an initial conditioning phase, building operational reliance and suppressing agent skepticism, before switching to an adversarial payload once an interaction threshold is reached. The evasion is temporal, not syntactic: benign at deploy time, th
Categories
Cite
@misc{rostamzadeh2026trustshiftprobe,
title = {{TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers}},
author = {Mehrdad Rostamzadeh and Sidhant Narula and Mohammad Ghasemigol and Daniel Takabi},
year = {2026},
month = aug,
eprint = {2608.23763},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.23763}
}