September 2026Unreviewed
SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs
Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer
Abstract
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA)
Categories
Cite
@misc{li2026safire,
title = {{SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs}},
author = {Pengfei Li and Naufal Suryanto and Sicheng Zhang and Mohammad Alsharid and Muzammal Naseer},
year = {2026},
month = sep,
eprint = {2609.07823},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.07823}
}