Skip to content
Search
paperJuly 2026Unreviewed

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

Yujiao Chen

Abstract

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{chen2026institutional,
  title = {{Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety}},
  author = {Yujiao Chen},
  year = {2026},
  month = jul,
  eprint = {2607.07695},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.07695}
}