Skip to content
Search
paperMarch 2026Unreviewed

IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs

Chuan Guo, J. Felipe, Cerón Uribe, Sicheng Zhu, Christopher A. Choquette-Choo, Stephanie L. Lin, Nikhil Kandpal, Milad Nasr, Michael Pokorny, S. Toyer, Miles Wang, Yao-Ching Yu, Alex Beutel, Kai Xiao OpenAI

Abstract

Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We i

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{guo2026ihchallenge,
  title = {{IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs}},
  author = {Chuan Guo and J. Felipe and Cerón Uribe and Sicheng Zhu and Christopher A. Choquette-Choo and Stephanie L. Lin and Nikhil Kandpal and Milad Nasr and Michael Pokorny and S. Toyer and Miles Wang and Yao-Ching Yu and Alex Beutel and {Kai Xiao OpenAI}},
  year = {2026},
  month = mar,
  eprint = {2603.10521},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/0d1a1ced045bbc21126efb2be0a4064e1ce18f63}
}