Skip to content
Search
paperAugust 2026Unreviewed

Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

S. Samarakoon, M. A. V. J. Muthugala, W. K. R. Sachinthana, M. R. Elara

Abstract

Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception and instruction-following introduces a new attack surface: adversarial text placed within the robot's visual field can act as an indirect prompt injection into the VLM's reasoning stack. We present a systematic study of physical prompt injection attacks against

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{samarakoon2026hijacking,
  title = {{Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots}},
  author = {S. Samarakoon and M. A. V. J. Muthugala and W. K. R. Sachinthana and M. R. Elara},
  year = {2026},
  month = aug,
  eprint = {2608.05715},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/3fa9e3ec5f9b01ff3e61b01227b43db042194334}
}