Skip to content
Search
paperJune 2026Unreviewed

Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs

Charbel El Feghali, Arkil Patel, Nicholas Meade, Spandana Gella, Verna Dankers, Siva Reddy

Abstract

Open-weight Large Language Models (LLMs) enable scientific progress and broad deployment. However, they make it difficult to control access to sensitive capabilities. Current practice either suppresses dangerous capabilities before release or mediates access through closed services that use specialized model variants, input/output monitors, and API permissions. The former is susceptible to jailbreaks while sacrificing capability for all users to mitigate the risks posed by a few, and the latter

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{feghali2026open,
  title = {{Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs}},
  author = {Charbel El Feghali and Arkil Patel and Nicholas Meade and Spandana Gella and Verna Dankers and Siva Reddy},
  year = {2026},
  month = jun,
  eprint = {2606.21638},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.21638}
}