Skip to content
Search
paperMay 2026Unreviewed

When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents

Strick Sheng, Ziyue Wang, Liyi Zhou

Abstract

Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations influence tool use, state tracking, and action sequencing, yet their reliability and authority are often uncertain. Environmental grounding is therefore a systems-level problem involving context admission, evidence provenance, freshness checking, verification policy, action gating, and model reasoning. Existing agent benchmarks mainly evaluate tas

Categories

Framework mappings

OWASP Top 10 for Agentic Applications
  • ASI02Tool Misuse & Exploitation
MITRE ATLAS
  • AML.T0053AI Agent Tool Invocation

Suggested from the entry's categories.

Cite

@misc{sheng2026when,
  title = {{When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents}},
  author = {Strick Sheng and Ziyue Wang and Liyi Zhou},
  year = {2026},
  month = may,
  eprint = {2605.08828},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.08828}
}