← Back to feed
2026-07-07visionmultimodalinfra

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Cong Su, Jiaju Han, Xuemeng Sun, Chengyin Hu, Qike Zhang, Jiujiang Guo, Yiwei Wei, Jiahuan Long

PDF preview for AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
Read on arXiv →

Key claim

AirflowAttack exposes vulnerabilities in infrared vision-language models.

In plain English

Vision-language models are increasingly used in security settings with infrared imagery, but their robustness against adversarial attacks is not well understood. Current methods do not address the unique challenges posed by infrared data. This paper introduces AirflowAttack, the first attack that uses thermal-airflow turbulence to create effective perturbations, achieving a high attack success rate across various models. Builders should care because it reveals critical vulnerabilities in the rapidly evolving field of infrared vision-language models, which could impact security applications.

Novelty
8.5/10

Introduces a novel adversarial attack method for IR VLMs using airflow turbulence.

Reliability
7.5/10

Demonstrates effectiveness against multiple models with solid benchmarking.

Deep reliability assessment

The methodology supports the claim that AirflowAttack can significantly reduce scene-classification accuracy in IR VLMs, but the physical realizability of the attack remains untested, which is a critical aspect for real-world applicability.

Reproducibility

No open source code or dataset is mentioned in the paper.

Key figure

Figure 1 provides an overview of AirflowAttack, illustrating how a lightweight generator maps a low-dimensional latent code to a thermal-airflow perturbation optimized on a surrogate IR-finetuned CLIP model.

Benchmark results

~five diverse CLIP backbonesattack success rate (ASR): 48.5vs four IR-specific physical baselines+10.8% to +20.8%SOTA