Skip to content
Search
paperSeptember 2026Unreviewed

Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

Zhongan Bi, Qiwen Wang, Jianrong Jiang, Jigang Ding, Wenwen Xiong, Changhua Meng, Xuanang Gao, Kepeng Lin, Changjiang Jiang, Yiang Chen, Huan Yao, Wei Wang, Zhenyu Ma, Wenhui Dong

Abstract

Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. Existing benchmarks largely measure whether manipulated content is retrieved or endorsed, but do not track whether an agent verifies suspicious evidence, revises adopted claims, or recovers before producing its final recommendation. We introduce HAE-GEO, a benchmark that tracks the full trajectory from exposure to recovery under progressively more pe

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{bi2026evaluating,
  title = {{Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning}},
  author = {Zhongan Bi and Qiwen Wang and Jianrong Jiang and Jigang Ding and Wenwen Xiong and Changhua Meng and Xuanang Gao and Kepeng Lin and Changjiang Jiang and Yiang Chen and Huan Yao and Wei Wang and Zhenyu Ma and Wenhui Dong},
  year = {2026},
  month = sep,
  eprint = {2609.06027},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06027}
}