Skip to content
Search
paperMarch 2026Unreviewed

Detection of adversarial intent in Human-AI teams using LLMs

Abed K. Musaffar, Ambuj Singh, Francesco Bullo

Abstract

Large language models (LLMs) are increasingly deployed in human-AI teams as support agents for complex tasks such as information retrieval, programming, and decision-making assistance. While these agents' autonomy and contextual knowledge enables them to be useful, it also exposes them to a broad range of attacks, including data poisoning, prompt injection, and even prompt engineering. Through these attack vectors, malicious actors can manipulate an LLM agent to provide harmful information, pote

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{musaffar2026detection,
  title = {{Detection of adversarial intent in Human-AI teams using LLMs}},
  author = {Abed K. Musaffar and Ambuj Singh and Francesco Bullo},
  year = {2026},
  month = mar,
  eprint = {2603.20976},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2603.20976}
}