Skip to content
Search
paperJuly 2026Unreviewed

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Brett Reynolds

Abstract

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judg

Categories

Cite

@misc{reynolds2026adversarial,
  title = {{Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity}},
  author = {Brett Reynolds},
  year = {2026},
  month = jul,
  eprint = {2607.01153},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.01153}
}