Skip to content
Search
paperSeptember 2026Unreviewed

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer

Abstract

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA)

Categories

Cite

@misc{li2026safire,
  title = {{SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs}},
  author = {Pengfei Li and Naufal Suryanto and Sicheng Zhang and Mohammad Alsharid and Muzammal Naseer},
  year = {2026},
  month = sep,
  eprint = {2609.07823},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.07823}
}