June 2026Unreviewed
MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG
Inderjeet Singh, Andrés Murillo, Motoyoshi Sekiya, Yuki Unno, Junichi Suga
Abstract
Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-teaming approaches are typically surface-specific and often recycle known attack templates; on text-poisoning benchmarks we measure 73-84% exact duplication. We present MIRROR, a unified cross-surface framework that performs memory-guided Monte Carlo tree search wh
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
- AML.T0051LLM Prompt Injection
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{singh2026mirror,
title = {{MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG}},
author = {Inderjeet Singh and Andrés Murillo and Motoyoshi Sekiya and Yuki Unno and Junichi Suga},
year = {2026},
month = jun,
eprint = {2606.26793},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.26793}
}