EMNLP 2026 Findings

SAFIRE SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Pengfei Li1 Naufal Suryanto1 Sicheng Zhang1 Mohammad Alsharid1 Muzammal Naseer1,2
1Khalifa University
2University of Western Australia
muhammadmuzammal.naseer@ku.ac.ae
Paper Code Dataset

Abstract

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning. A GPT-5.4-assisted multi-stage verification pipeline with MLLM majority voting ensures annotation quality. Evaluating ten open-source MLLMs (8B-38B) yields an average accuracy of 61.9%, exposing major gaps in safety-critical reasoning. We further show that adapting vision encoders with only 7% of our domain-specific data boosts fire-scene classification accuracy from 20.1% to 64.5%, indicating that carefully curated data can yield substantial gains even when data volume is limited.

A Comprehensive Benchmark

SAFIRE contains 82,842 captioned images across 20 fire-smoke scenarios, with 9,668 MCVQA-active images generating 192,904 question-answer pairs for fine-grained multimodal reasoning.

82,842
Captioned Images
192,904
MCVQA Pairs
9,668
QA-Active Images
20
Scenarios
10
Reasoning Dimensions

5 Process Groups

  • 🌋Natural Phenomena Open Grassland, Forest Fire, Volcano, Meteor
  • 🏭Industrial Operations Aerospace Operation, Flare Stack, Metal Forging
  • 💥Accidental Incidents Residential Fire, Explosion, Vehicle Fire
  • 🎆Recreational Activities Firework, Sky Lantern-Hot Balloon, Barbecue, Campfire-Bonfire, Torch Parade
  • 🕯️Civil Controlled Scenes Waste Disposal, Gas Stove, Incense Burning, Candlelight, Smoking

10 Evaluation Dimensions

🔢 Target Counting
🗂️ Classification
🧠 General Reasoning
🎯 Fire/Smoke Intention
🎭 Emotional Response
💬 Polysemy
🔍 Fire/Smoke Attributes
🗺️ Fire/Smoke Spatial Correlation
📍 Position Identification
🧍 Human Presence

The SAFIRE Pipeline

SAFIRE follows four stages: multi-source collection and filtering, MCVQA/caption construction with hierarchical labels, MLLM evaluation across ten reasoning dimensions, and vision-language encoder benchmarking under zero-shot and few-shot settings.

SAFIRE Pipeline Diagram

Illustrative MCVQA Samples

The 12 examples below follow the paper's sample set and span scene understanding, target counting, classification, positional information, spatial relation, danger level, broader indication, reasoning, fire intention, fire environment, prediction, and language polysemy.

Scene Understanding (Open Grassland)
Open Grassland fire
Image Placeholder
(Open Grassland)

What kind of flame is this?

  • A. Industrial fire,
  • B. Aerial fireworks burst,
  • C. Home accident fire,
  • D. Crop-residue fire
Target Counting (Barbecue)
Barbecue
Image Placeholder
(Barbecue)

How many distinct fire are visible?

  • A. None,
  • B. One,
  • C. Two,
  • D. More than two
Classification (Campfire-Bonfire)
Campfire-Bonfire
Image Placeholder
(Campfire-Bonfire)

What kind of smoke/flare?

  • A. Campfire or fire pit
  • B. Grill or cooking fire,
  • C. Not Applicable,
  • D. Icon or symbolic fire
Positional Info. (Aerospace Operation)
Aerospace Operation
Image Placeholder
(Aerospace Operation)

Where is the fire located?

  • A. Left side,
  • B. Right side,
  • C. Centre,
  • D. Not visible
Spatial Relation (Firework)
Firework
Image Placeholder
(Firework)

What is the spatial relation?

  • A. Smoke above,
  • B. Smoke to the left,
  • C. Smoke to the right,
  • D. Not applicable
Danger Level (Volcano)
Volcano
Image Placeholder
(Volcano)

What is the apparent danger level?

  • A. Low,
  • B. Moderate,
  • C. High
  • D. N/A
Broader Indication (Flare Stack)
Flare Stack
Image Placeholder
(Flare Stack)

Is the fire controlled?

  • A. Yes,
  • B. No,
  • C. No fire,
  • D. Cannot tell
Reasoning (Vehicle Fire)
Vehicle Fire
Image Placeholder
(Vehicle Fire)

How would people react to this fire/smoke?

  • A. Call for help,
  • B. Ignore it,
  • C. Evacuate the area,
  • D. Take pictures
Fire Intention (Torch)
Torch
Image Placeholder
(Torch)

What is the purpose of this fire?

  • A. Accident,
  • B. Entertainment,
  • C. Industry production,
  • D. Culture activity
Fire Environment (Candlelight)
Candlelight
Image Placeholder
(Candlelight)

What time of the day does this fire suggests?

  • A. Day time,
  • B. Nighttime,
  • C. Indoor lighting,
  • D. Not determinable
Prediction (Incense Burning)
Incense Burning
Image Placeholder
(Incense Burning)

What will happen next to the Fire/Smoke?

  • A. Fire expands,
  • B. Smoke disperses,
  • C. Fire quenches,
  • D. Explosion occurs
Language Polysemy (Meteorite Fire)
Meteorite Fire
Image Placeholder
(Meteorite Fire)

What is "on fire" means here?

  • A. Literal Fire,
  • B. Excited situation,
  • C. Under attack,
  • D. Figurative speech

MLLM Evaluation Results

Ten open-source MLLMs from 8B to 38B achieve 61.87% average accuracy on the full SAFIRE MCVQA benchmark, showing that safety-critical fire-smoke reasoning remains difficult even for recent vision-language models.

Table 3: Accuracy (%) of MLLMs on the 20-scenario fire-smoke MCVQA benchmark.

Bold indicates the highest accuracy in each column, underlined indicates the second-highest accuracy, and † denotes the thinking-enabled model version. Avg. is the overall average; Scene Avg. is the macro average.

Model Size (B) Natural Phenomena Industrial Operations Accidental Incidents
Grassland Volcano Forest Meteor Aerospace Flare Stack Forging Residential Explos. Vehicle
InternVL3.5 8 54.13 56.96 58.85 50.60 53.13 60.85 55.48 58.35 53.06 59.46
GLM4.1V 9 59.58 57.85 60.13 57.66 55.56 75.37 57.21 60.87 56.18 57.20
Qwen3.5 9 68.97 58.45 63.97 68.28 62.37 73.14 60.81 65.17 60.73 63.03
Qwen3.5† 9 64.30 57.50 61.43 59.09 59.46 74.50 56.06 61.52 56.58 60.66
Gemma-3 12 58.42 55.66 56.91 53.29 53.09 60.68 52.08 57.85 56.54 55.31
Gemma-3 27 58.24 55.49 60.23 53.67 52.11 66.17 54.35 57.98 54.25 58.57
Qwen3-VL 32 65.70 61.67 63.36 60.04 60.02 74.94 62.29 64.00 55.97 65.66
Qwen3-VL† 32 69.26 60.29 63.75 61.41 61.55 72.32 59.77 63.77 58.55 64.57
Qwen3.6 35A3 69.32 61.60 63.15 62.17 61.65 74.29 61.24 64.17 57.16 64.69
InternVL3.5 38 66.39 58.64 61.82 59.41 58.04 73.30 61.42 62.19 55.61 61.77
Scene Avg. 63.43 58.41 61.36 58.56 57.70 70.56 58.07 61.59 56.46 61.09
Model Size (B) Recreational Activities Civil Controlled Scenes All
Avg.
Firework SkyLantern Barbecue Campfire Torch Waste Stove Incense Candle Smoking
InternVL3.5 8 54.39 62.14 55.75 65.75 70.27 60.91 67.87 46.82 65.86 60.34 58.37
GLM4.1V 9 52.09 57.85 50.26 60.08 66.85 59.63 70.51 58.13 69.39 60.68 59.70
Qwen3.5 9 62.18 70.38 57.55 69.93 75.02 61.03 77.34 57.29 70.96 64.13 64.89
Qwen3.5† 9 59.72 67.10 55.02 69.12 74.84 58.87 73.83 53.19 67.85 61.20 62.17
Gemma-3 12 56.44 56.52 48.56 59.11 69.42 58.29 62.30 50.18 56.80 50.60 56.31
Gemma-3 27 58.25 62.53 50.42 65.01 73.57 60.04 65.20 45.12 61.82 55.06 58.10
Qwen3-VL 32 63.66 70.54 61.10 70.02 73.87 63.97 73.64 61.31 74.65 68.26 65.35
Qwen3-VL† 32 63.29 73.07 61.23 69.48 76.44 65.73 75.07 62.14 73.35 66.28 65.67
Qwen3.6 35A3 63.64 71.56 60.98 70.51 76.92 62.71 75.92 58.77 71.23 67.50 65.53
InternVL3.5 38 54.55 68.02 57.99 68.85 73.96 63.71 70.67 54.48 73.28 59.15 62.65
Scene Avg. 58.82 65.97 55.89 66.79 73.12 61.49 71.23 54.74 68.52 61.32

Table 4: Accuracy (%) across 10 evaluation dimensions on SAFIRE.

Bold indicates the highest accuracy per dimension, underlined indicates the second-highest accuracy, and † denotes the thinking-enabled model version.

Model Target
Counting
Class. General
Reasoning
Fire/Smoke
Intention
Emotional
Response
Polysemy Fire/Smoke
Attributes
Spatial
Corr.
Position
Identification
Human
Presence
InternVL3.5 (8B) 67.40 73.12 49.84 67.75 54.73 59.56 58.95 34.18 55.63 88.20
GLM4.1V (9B) 70.12 74.52 48.87 68.32 54.24 60.60 63.02 29.38 60.95 92.85
Qwen3.5 (9B) 66.41 77.29 55.81 68.50 70.69 70.68 65.80 48.83 61.13 93.07
Qwen3.5† (9B) 67.78 76.38 52.32 66.85 66.34 68.13 63.76 45.57 52.14 91.06
Gemma-3 (12B) 42.21 71.75 49.76 71.43 45.24 66.39 56.13 38.08 54.98 85.09
Gemma-3 (27B) 48.87 76.05 50.66 69.82 49.02 63.06 57.73 42.48 56.58 87.71
Qwen3-VL (32B) 67.34 79.20 55.89 73.76 69.63 71.61 65.98 48.23 60.59 92.15
Qwen3-VL† (32B) 66.78 81.85 57.58 69.64 69.79 64.54 66.54 53.77 58.50 90.49
Qwen3.6 (35B/A3B) 70.70 76.30 57.87 70.82 69.90 65.04 66.38 49.59 61.16 91.62
InternVL3.5 (38B) 63.49 77.40 54.38 71.93 67.39 66.51 63.76 38.90 54.32 90.58
Avg.(%) 63.11 76.39 53.30 69.88 61.70 65.61 62.80 42.90 57.60 90.28

Citation

If you find SAFIRE useful for your research, please consider citing our paper.

@inproceedings{li2026safire,
  title={SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs},
  author={Li, Pengfei and Suryanto, Naufal and Zhang, Sicheng and Alsharid, Mohammad and Naseer, Muzammal},
  booktitle={The 2026 Conference on Empirical Methods in Natural Language Processing},
  year={2026}
}