June 2026Unreviewed
Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs
Charbel El Feghali, Arkil Patel, Nicholas Meade, Spandana Gella, Verna Dankers, Siva Reddy
Abstract
Open-weight Large Language Models (LLMs) enable scientific progress and broad deployment. However, they make it difficult to control access to sensitive capabilities. Current practice either suppresses dangerous capabilities before release or mediates access through closed services that use specialized model variants, input/output monitors, and API permissions. The former is susceptible to jailbreaks while sacrificing capability for all users to mitigate the risks posed by a few, and the latter
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{feghali2026open,
title = {{Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs}},
author = {Charbel El Feghali and Arkil Patel and Nicholas Meade and Spandana Gella and Verna Dankers and Siva Reddy},
year = {2026},
month = jun,
eprint = {2606.21638},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.21638}
}