Skip to content
Search
paperJune 2026Unreviewed

Training LLMs to Enforce Multi-Level Instruction Hierarchies via Gravity-Weighted Direct Preference Optimization

Lena S. Bolliger, Lena A. Jäger

Abstract

Production LLMs receive instructions from sources with very different levels of trust, yet attend to every token with uniform architectural privilege. This is the structural vulnerability that enables malicious prompt injections and, more broadly, leaves models without a principled way to resolve conflicts between legitimate but competing instructions. A common training-based response is to teach models an explicit instruction hierarchy; existing approaches, however, formalize hierarchies of onl

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection

Suggested from the entry's categories.

Cite

@misc{bolliger2026training,
  title = {{Training LLMs to Enforce Multi-Level Instruction Hierarchies via Gravity-Weighted Direct Preference Optimization}},
  author = {Lena S. Bolliger and Lena A. Jäger},
  year = {2026},
  month = jun,
  eprint = {2606.10860},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.10860}
}