Skip to content
Search
paperMay 2026Unreviewed

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

Aditya Nawal, Manit Baser, Mohan Gurusamy

Abstract

AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the generation pipeline can weaken the safety alignment mechanisms that govern model outputs. Prior work shows that enabling retrieval in agents increases compliance with harmful requests. We introduce AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents. The framework examines t

Categories

Cite

@misc{nawal2026relevance,
  title = {{Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents}},
  author = {Aditya Nawal and Manit Baser and Mohan Gurusamy},
  year = {2026},
  month = may,
  eprint = {2605.29224},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.29224}
}