September 2026Unreviewed
TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors
Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho
Abstract
Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompts with varying threat implicitness. We introduce TIER, a Threat Implicitness Benchmark for behavioral safety evaluation of LLMs. TIER covers four risk domains and four threat levels, from explicit harmful requests to sophisticated jailbreaks. Responses are assessed using a six-label behavior scale and two independent LLM judges. Experiments on six open-weight LLMs show that safety behavi
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{trinhthi2026tier,
title = {{TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors}},
author = {Thu-Hien Trinh-Thi and Hai-Yen Vong and Thanh-Ha Ung-Dung and Tram Ho},
year = {2026},
month = sep,
eprint = {2609.05117},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.05117}
}