Skip to content
Search
paperJuly 2026Unreviewed

Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

Mingxiao Liu, Yitong Li, Haoren Zhao, Yaoxiang Bian, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang

Abstract

Large Language Model (LLM)-driven multimodal agents are increasingly deployed to execute autonomous tasks via continuous audio interaction. While this paradigm enhances interaction naturalness, it introduces a critical yet under-explored attack surface, as audio inputs inevitably contain environmental noise beyond user control. In this paper, we investigate concurrent audio prompt injection attacks targeting multimodal agents. Distinct from traditional acoustic attacks on voice devices, we propo

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{liu2026piggybacking,
  title = {{Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents}},
  author = {Mingxiao Liu and Yitong Li and Haoren Zhao and Yaoxiang Bian and Jianan Ma and Jian Zhang and Jialuo Chen and Xinhao Deng and Zhen Wang},
  year = {2026},
  month = jul,
  eprint = {2607.28165},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.28165}
}