:
Safety-Critical Benchmark for Fine-grained Fire and Smoke
Understanding in Multimodal LLMs
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning. A GPT-5.4-assisted multi-stage verification pipeline with MLLM majority voting ensures annotation quality. Evaluating ten open-source MLLMs (8B-38B) yields an average accuracy of 61.9%, exposing major gaps in safety-critical reasoning. We further show that adapting vision encoders with only 7% of our domain-specific data boosts fire-scene classification accuracy from 20.1% to 64.5%, indicating that carefully curated data can yield substantial gains even when data volume is limited.
SAFIRE contains 82,842 captioned images across 20 fire-smoke scenarios, with 9,668 MCVQA-active images generating 192,904 question-answer pairs for fine-grained multimodal reasoning.
SAFIRE follows four stages: multi-source collection and filtering, MCVQA/caption construction with hierarchical labels, MLLM evaluation across ten reasoning dimensions, and vision-language encoder benchmarking under zero-shot and few-shot settings.
The 12 examples below follow the paper's sample set and span scene understanding, target counting, classification, positional information, spatial relation, danger level, broader indication, reasoning, fire intention, fire environment, prediction, and language polysemy.
What kind of flame is this?
How many distinct fire are visible?
What kind of smoke/flare?
Where is the fire located?
What is the spatial relation?
What is the apparent danger level?
Is the fire controlled?
How would people react to this fire/smoke?
What is the purpose of this fire?
What time of the day does this fire suggests?
What will happen next to the Fire/Smoke?
What is "on fire" means here?
Ten open-source MLLMs from 8B to 38B achieve 61.87% average accuracy on the full SAFIRE MCVQA benchmark, showing that safety-critical fire-smoke reasoning remains difficult even for recent vision-language models.
Bold indicates the highest accuracy in each column, underlined indicates the second-highest accuracy, and † denotes the thinking-enabled model version. Avg. is the overall average; Scene Avg. is the macro average.
| Model | Size (B) | Natural Phenomena | Industrial Operations | Accidental Incidents | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Grassland | Volcano | Forest | Meteor | Aerospace | Flare Stack | Forging | Residential | Explos. | Vehicle | ||
| InternVL3.5 | 8 | 54.13 | 56.96 | 58.85 | 50.60 | 53.13 | 60.85 | 55.48 | 58.35 | 53.06 | 59.46 |
| GLM4.1V | 9 | 59.58 | 57.85 | 60.13 | 57.66 | 55.56 | 75.37 | 57.21 | 60.87 | 56.18 | 57.20 |
| Qwen3.5 | 9 | 68.97 | 58.45 | 63.97 | 68.28 | 62.37 | 73.14 | 60.81 | 65.17 | 60.73 | 63.03 |
| Qwen3.5† | 9 | 64.30 | 57.50 | 61.43 | 59.09 | 59.46 | 74.50 | 56.06 | 61.52 | 56.58 | 60.66 |
| Gemma-3 | 12 | 58.42 | 55.66 | 56.91 | 53.29 | 53.09 | 60.68 | 52.08 | 57.85 | 56.54 | 55.31 |
| Gemma-3 | 27 | 58.24 | 55.49 | 60.23 | 53.67 | 52.11 | 66.17 | 54.35 | 57.98 | 54.25 | 58.57 |
| Qwen3-VL | 32 | 65.70 | 61.67 | 63.36 | 60.04 | 60.02 | 74.94 | 62.29 | 64.00 | 55.97 | 65.66 |
| Qwen3-VL† | 32 | 69.26 | 60.29 | 63.75 | 61.41 | 61.55 | 72.32 | 59.77 | 63.77 | 58.55 | 64.57 |
| Qwen3.6 | 35A3 | 69.32 | 61.60 | 63.15 | 62.17 | 61.65 | 74.29 | 61.24 | 64.17 | 57.16 | 64.69 |
| InternVL3.5 | 38 | 66.39 | 58.64 | 61.82 | 59.41 | 58.04 | 73.30 | 61.42 | 62.19 | 55.61 | 61.77 |
| Scene Avg. | – | 63.43 | 58.41 | 61.36 | 58.56 | 57.70 | 70.56 | 58.07 | 61.59 | 56.46 | 61.09 |
| Model | Size (B) | Recreational Activities | Civil Controlled Scenes | All Avg. |
||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Firework | SkyLantern | Barbecue | Campfire | Torch | Waste | Stove | Incense | Candle | Smoking | |||
| InternVL3.5 | 8 | 54.39 | 62.14 | 55.75 | 65.75 | 70.27 | 60.91 | 67.87 | 46.82 | 65.86 | 60.34 | 58.37 |
| GLM4.1V | 9 | 52.09 | 57.85 | 50.26 | 60.08 | 66.85 | 59.63 | 70.51 | 58.13 | 69.39 | 60.68 | 59.70 |
| Qwen3.5 | 9 | 62.18 | 70.38 | 57.55 | 69.93 | 75.02 | 61.03 | 77.34 | 57.29 | 70.96 | 64.13 | 64.89 |
| Qwen3.5† | 9 | 59.72 | 67.10 | 55.02 | 69.12 | 74.84 | 58.87 | 73.83 | 53.19 | 67.85 | 61.20 | 62.17 |
| Gemma-3 | 12 | 56.44 | 56.52 | 48.56 | 59.11 | 69.42 | 58.29 | 62.30 | 50.18 | 56.80 | 50.60 | 56.31 |
| Gemma-3 | 27 | 58.25 | 62.53 | 50.42 | 65.01 | 73.57 | 60.04 | 65.20 | 45.12 | 61.82 | 55.06 | 58.10 |
| Qwen3-VL | 32 | 63.66 | 70.54 | 61.10 | 70.02 | 73.87 | 63.97 | 73.64 | 61.31 | 74.65 | 68.26 | 65.35 |
| Qwen3-VL† | 32 | 63.29 | 73.07 | 61.23 | 69.48 | 76.44 | 65.73 | 75.07 | 62.14 | 73.35 | 66.28 | 65.67 |
| Qwen3.6 | 35A3 | 63.64 | 71.56 | 60.98 | 70.51 | 76.92 | 62.71 | 75.92 | 58.77 | 71.23 | 67.50 | 65.53 |
| InternVL3.5 | 38 | 54.55 | 68.02 | 57.99 | 68.85 | 73.96 | 63.71 | 70.67 | 54.48 | 73.28 | 59.15 | 62.65 |
| Scene Avg. | – | 58.82 | 65.97 | 55.89 | 66.79 | 73.12 | 61.49 | 71.23 | 54.74 | 68.52 | 61.32 | – |
Bold indicates the highest accuracy per dimension, underlined indicates the second-highest accuracy, and † denotes the thinking-enabled model version.
| Model | Target Counting |
Class. | General Reasoning |
Fire/Smoke Intention |
Emotional Response |
Polysemy | Fire/Smoke Attributes |
Spatial Corr. |
Position Identification |
Human Presence |
|---|---|---|---|---|---|---|---|---|---|---|
| InternVL3.5 (8B) | 67.40 | 73.12 | 49.84 | 67.75 | 54.73 | 59.56 | 58.95 | 34.18 | 55.63 | 88.20 |
| GLM4.1V (9B) | 70.12 | 74.52 | 48.87 | 68.32 | 54.24 | 60.60 | 63.02 | 29.38 | 60.95 | 92.85 |
| Qwen3.5 (9B) | 66.41 | 77.29 | 55.81 | 68.50 | 70.69 | 70.68 | 65.80 | 48.83 | 61.13 | 93.07 |
| Qwen3.5† (9B) | 67.78 | 76.38 | 52.32 | 66.85 | 66.34 | 68.13 | 63.76 | 45.57 | 52.14 | 91.06 |
| Gemma-3 (12B) | 42.21 | 71.75 | 49.76 | 71.43 | 45.24 | 66.39 | 56.13 | 38.08 | 54.98 | 85.09 |
| Gemma-3 (27B) | 48.87 | 76.05 | 50.66 | 69.82 | 49.02 | 63.06 | 57.73 | 42.48 | 56.58 | 87.71 |
| Qwen3-VL (32B) | 67.34 | 79.20 | 55.89 | 73.76 | 69.63 | 71.61 | 65.98 | 48.23 | 60.59 | 92.15 |
| Qwen3-VL† (32B) | 66.78 | 81.85 | 57.58 | 69.64 | 69.79 | 64.54 | 66.54 | 53.77 | 58.50 | 90.49 |
| Qwen3.6 (35B/A3B) | 70.70 | 76.30 | 57.87 | 70.82 | 69.90 | 65.04 | 66.38 | 49.59 | 61.16 | 91.62 |
| InternVL3.5 (38B) | 63.49 | 77.40 | 54.38 | 71.93 | 67.39 | 66.51 | 63.76 | 38.90 | 54.32 | 90.58 |
| Avg.(%) | 63.11 | 76.39 | 53.30 | 69.88 | 61.70 | 65.61 | 62.80 | 42.90 | 57.60 | 90.28 |
If you find SAFIRE useful for your research, please consider citing our paper.
@inproceedings{li2026safire,
title={SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs},
author={Li, Pengfei and Suryanto, Naufal and Zhang, Sicheng and Alsharid, Mohammad and Naseer, Muzammal},
booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing},
year={2026}
}