Skip to content
Search
paperAugust 2026Unreviewed

PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents

He Zhang, Fei-Long Li, Ding-Ning Long, Yi Cui, Peijun Zhang, Yue-Wen Zhang, Qianyao Xu, Xinyi Fu

Abstract

Smart-home assistants increasingly use multimodal large language models (MLLMs) that perceive video and audio directly. This raises a safety question specific to the home: can the agent tell a genuine user command from ambient or externally-sourced content, television speech, on-screen text, or an overheard conversation, that merely looks like a command? We introduce PromptShield-Home, a pilot benchmark of realistic smart-home scenarios spanning addressee ambiguity, screen/audio injection, healt

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{zhang2026promptshield,
  title = {{PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents}},
  author = {He Zhang and Fei-Long Li and Ding-Ning Long and Yi Cui and Peijun Zhang and Yue-Wen Zhang and Qianyao Xu and Xinyi Fu},
  year = {2026},
  month = aug,
  eprint = {2608.05495},
  archivePrefix = {arXiv},
  doi = {10.1145/3798063.3837198},
  url = {https://www.semanticscholar.org/paper/30d00b2862a54156f390d455743bd002230cb3bb}
}