July 2026Unreviewed
Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels
Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun
Abstract
Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are sensitive to the benchmark's grading procedure and capture only observed behavior on a given set of attacks, without directly revealing the hidden fragility of the underlying safety mechanisms. This work proposes JADR (Jacobian Assessment of Danger Recognition), a protocol that measures a model's internal representation through Jacobian space (J-spac
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{prosvirnin2026silent,
title = {{Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels}},
author = {Roman Prosvirnin and Victor Minchenkov and Alexey Soldatov and Vladimir Bashun},
year = {2026},
month = jul,
eprint = {2607.12792},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.12792}
}