Skip to content
Search
paperApril 2026Unreviewed

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection

Meng Chen, Kun Wang, Li Lu, Jiaheng Zhang, Tianwei Zhang

Abstract

Modern Large audio-language models (LALMs) power intelligent voice interactions by tightly integrating audio and text. This integration, however, expands the attack surface beyond text and introduces vulnerabilities in the continuous, high-dimensional audio channel. While prior work studied audio jailbreaks, the security risks of malicious audio injection and downstream behavior manipulation remain underexamined. In this work, we reveal a previously overlooked threat, auditory prompt injection,

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chen2026hijacking,
  title = {{Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection}},
  author = {Meng Chen and Kun Wang and Li Lu and Jiaheng Zhang and Tianwei Zhang},
  year = {2026},
  month = apr,
  eprint = {2604.14604},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.14604}
}