Semantic Segmentation Model for RGB-T Images with Cross-modal Interaction and Dynamic Fusion
HE Xinyi
ZHU Li
TAN Hanzhong
LI Yun
ZHOU Yuanxin
Abstract:Aiming at the accuracy degradation problem caused by modal feature mismatch and thermal imaging overexposure noise in the semantic segmentation of red-green-blue-thermal(RGB-T)images,a cross-modal interaction and dynamic fusion for RGB-T image semantic segmentation(CIDF)model was proposed to realize efficient co-sensing of RGB-T modalities.Firstly,a cross-modal skip connection(CMSC)mechanism was designed to establish multi-scale feature interactions between RGB-T modalities at the encoding stage,and effectively mitigate the inter-modal feature distribution differences through cross-layer feature transfer and adaptive aggregation strategies.Secondly,the multi-modal convolution fusion module(MCFM)was proposed,where dynamic feature preference was achieved through convolution,and the overexposure noise interference in the thermal imaging modality was suppressed.The experiments were based on the publicly available multi-spectral image fusion net for urban road scenes(MFNet)dataset and a self-constructed dataset of photovoltaic inspection scenes.The results showed that the average accuracy of CIDF model reached 72.9%and the mean intersection over union of CIDF model reached 56.6%on the MFNet dataset.In addition,average accuracy and mean intersection over union on the photovoltaic inspection dataset reached 98.9%and 97.4%,respectively.The research results not only confirmed the effectiveness of the CIDF model,but also highlighted its superior performance in cross-scenario applications.
Keywords:CMSC moduleMCFM modulemultimodal fusiondual branch encoderRGB-T semantic segmentationscene adaptive
Publication Date:2025-09-20
Online Publishing Date:2025-09-24(First online date of this platform, not the publication date of the document)
Pages:6( 370-375 )