April 2026Unreviewed
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
Xuanli He, Bilgehan Sel, Faizan Ali, Jenny Bao, Hoagy Cunningham, Jerry Wei
Abstract
Large Language Models (LLMs) are increasingly exposed to adaptive jailbreaking, particularly in high-stakes Chemical, Biological, Radiological, and Nuclear (CBRN) domains. Although streaming probes enable real-time monitoring, they still make systematic errors. We identify a core issue: existing methods often rely on a few high-scoring tokens, leading to false alarms when sensitive CBRN terms appear in benign contexts. To address this, we introduce a streaming probing objective that requires mul
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{he2026segmentlevel,
title = {{Segment-Level Coherence for Robust Harmful Intent Probing in LLMs}},
author = {Xuanli He and Bilgehan Sel and Faizan Ali and Jenny Bao and Hoagy Cunningham and Jerry Wei},
year = {2026},
month = apr,
eprint = {2604.14865},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.14865}
}