Skip to content
Search
paperSeptember 2026Unreviewed

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho

Abstract

Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompts with varying threat implicitness. We introduce TIER, a Threat Implicitness Benchmark for behavioral safety evaluation of LLMs. TIER covers four risk domains and four threat levels, from explicit harmful requests to sophisticated jailbreaks. Responses are assessed using a six-label behavior scale and two independent LLM judges. Experiments on six open-weight LLMs show that safety behavi

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{trinhthi2026tier,
  title = {{TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors}},
  author = {Thu-Hien Trinh-Thi and Hai-Yen Vong and Thanh-Ha Ung-Dung and Tram Ho},
  year = {2026},
  month = sep,
  eprint = {2609.05117},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.05117}
}