March 2026Unreviewed
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Chuan Guo, J. Felipe, Cerón Uribe, Sicheng Zhu, Christopher A. Choquette-Choo, Stephanie L. Lin, Nikhil Kandpal, Milad Nasr, Michael Pokorny, S. Toyer, Miles Wang, Yao-Ching Yu, Alex Beutel, Kai Xiao OpenAI
Abstract
Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instruction conflicts. IH is key to defending against jailbreaks, system prompt extractions, and agentic prompt injections. However, robust IH behavior is difficult to train: IH failures can be confounded with instruction-following failures, conflicts can be nuanced, and models can learn shortcuts such as overrefusing. We i
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{guo2026ihchallenge,
title = {{IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs}},
author = {Chuan Guo and J. Felipe and Cerón Uribe and Sicheng Zhu and Christopher A. Choquette-Choo and Stephanie L. Lin and Nikhil Kandpal and Milad Nasr and Michael Pokorny and S. Toyer and Miles Wang and Yao-Ching Yu and Alex Beutel and {Kai Xiao OpenAI}},
year = {2026},
month = mar,
eprint = {2603.10521},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/0d1a1ced045bbc21126efb2be0a4064e1ce18f63}
}