Skip to content
Search
paperApril 2026Unreviewed

Segment-Level Coherence for Robust Harmful Intent Probing in LLMs

Xuanli He, Bilgehan Sel, Faizan Ali, Jenny Bao, Hoagy Cunningham, Jerry Wei

Abstract

Large Language Models (LLMs) are increasingly exposed to adaptive jailbreaking, particularly in high-stakes Chemical, Biological, Radiological, and Nuclear (CBRN) domains. Although streaming probes enable real-time monitoring, they still make systematic errors. We identify a core issue: existing methods often rely on a few high-scoring tokens, leading to false alarms when sensitive CBRN terms appear in benign contexts. To address this, we introduce a streaming probing objective that requires mul

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{he2026segmentlevel,
  title = {{Segment-Level Coherence for Robust Harmful Intent Probing in LLMs}},
  author = {Xuanli He and Bilgehan Sel and Faizan Ali and Jenny Bao and Hoagy Cunningham and Jerry Wei},
  year = {2026},
  month = apr,
  eprint = {2604.14865},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.14865}
}