August 2026Unreviewed
Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
S. Samarakoon, M. A. V. J. Muthugala, W. K. R. Sachinthana, M. R. Elara
Abstract
Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception and instruction-following introduces a new attack surface: adversarial text placed within the robot's visual field can act as an indirect prompt injection into the VLM's reasoning stack. We present a systematic study of physical prompt injection attacks against
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Suggested from the entry's categories.
Cite
@misc{samarakoon2026hijacking,
title = {{Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots}},
author = {S. Samarakoon and M. A. V. J. Muthugala and W. K. R. Sachinthana and M. R. Elara},
year = {2026},
month = aug,
eprint = {2608.05715},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/3fa9e3ec5f9b01ff3e61b01227b43db042194334}
}