HAPNet: Toward superior RGB-thermal scene parsing via hybrid, asymmetric, and progressive heterogeneous feature fusion

Jiahang Li , Peng Yun , Yang Xu , Ye Zhang , Mingjian Sun , Qijun Chen , Ilin Alexander , Rui Fan

Biomimetic Intelligence and Robotics ›› 2026, Vol. 6 ›› Issue (3) : 100309

PDF (3064KB)
Biomimetic Intelligence and Robotics ›› 2026, Vol. 6 ›› Issue (3) :100309 DOI: 10.1016/j.birob.2026.100309
Research Article
research-article
HAPNet: Toward superior RGB-thermal scene parsing via hybrid, asymmetric, and progressive heterogeneous feature fusion
Author information +
History +
PDF (3064KB)

Abstract

Data-fusion networks have shown significant promise for RGB-thermal scene parsing. However, the majority of existing studies have relied on symmetric duplex encoders for heterogeneous feature extraction and fusion, paying inadequate attention to the inherent differences between RGB and thermal modalities. Recent progress in vision foundation models (VFMs), which leverage self-supervised learning on large-scale unlabeled datasets, has exhibited superior capabilities in extracting informative, general-purpose features compared to supervised encoders. However, their potential has yet to be fully leveraged in the domain. In this study, we take one step toward this new research area by exploring a feasible strategy to fully exploit VFM features for RGB-thermal scene parsing. Specifically, we delve deeper into the unique characteristics of RGB and thermal modalities, thereby designing a hybrid, asymmetric encoder that incorporates both a VFM and a cross-modal spatial prior descriptor (CSPD), enabling enhanced extraction of complementary heterogeneous features. The extracted features undergo dual-path feature fusion through our proposed progressive heterogeneous feature integrators. Moreover, we introduce an auxiliary task to further enrich the local semantics of fused features, thereby improving the overall performance of RGB-thermal scene parsing. Our proposed HAPNet, incorporating all these components, delivers superior performance under challenging illumination conditions. Extensive experiments demonstrate that HAPNet outperforms all other state-of-the-art methods, with improvements of 0.1%, 1.0%, and 2.4% in mIoU on three public RGB-thermal scene parsing datasets: MFNet, PST900, and KP Day-Night, respectively. Additionally, our method exhibits exceptional generalizability for RGB-HHA scene parsing. We believe this new paradigm has opened up new opportunities for future developments in data-fusion scene parsing approaches. The source code is publicly available at https://mias.group/HAPNet/.

Keywords

Data-fusion / Thermal / Scene parsing / Heterogeneous feature / Vision foundation model

Cite this article

Download citation ▾
Jiahang Li, Peng Yun, Yang Xu, Ye Zhang, Mingjian Sun, Qijun Chen, Ilin Alexander, Rui Fan. HAPNet: Toward superior RGB-thermal scene parsing via hybrid, asymmetric, and progressive heterogeneous feature fusion. Biomimetic Intelligence and Robotics, 2026, 6 (3) : 100309 DOI:10.1016/j.birob.2026.100309

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

S.S. Shivakumar, et al., PST900: RGB-Thermal calibration, dataset and segmentation network, 2020 IEEE International Conference on Robotics and Automation, ICRA, IEEE, 2020, pp. 9441-9447.

[2]

M. Liang, et al., Explicit attention-enhanced fusion for RGB-Thermal perception tasks, IEEE Robot. Autom. Lett. 8 (7) (2023) 4060-4067.

[3]

Y. Sun, et al., RTFNet: RGB-Thermal fusion network for semantic segmentation of urban scenes, IEEE Robot. Autom. Lett. 4 (3) (2019) 2576-2583.

[4]

Q. Ha, et al., MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes, 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS, IEEE, 2017, pp. 5108-5115.

[5]

Q. Zhang, et al., ABMDRNet: Adaptive-weighted Bi-directional modality difference reduction network for RGB-T semantic segmentation, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2021, pp. 2633-2642.

[6]

S. Guo, et al., LIX: Implicitly infusing spatial geometric prior knowledge into visual semantic segmentation for autonomous driving, IEEE Trans. Image Process. 34 (2025) 7250-7263, https://doi.org/10.1109/TIP.2025.3618378.

[7]

U. Shin, et al., Complementary random masking for RGB-Thermal semantic segmentation, 2024 IEEE International Conference on Robotics and Automation, ICRA, IEEE, 2024, pp. 11110-11117.

[8]

W. Zhou, et al., GMNet: Graded-feature multilabel-learning network for RGB-Thermal urban scene semantic segmentation, IEEE Trans. Image Process. 30 (2021) 7790-7802.

[9]

Y. Lv, et al., Context-aware interaction network for RGB-T semantic segmentation, IEEE Trans. Multimed., (2024),DOI:10.1109/TMM.2023.3349072.

[10]

Z. Wu, et al., S3M-Net: Joint learning of semantic segmentation and stereo matching for autonomous driving, IEEE Trans. Intell. Veh. 9 (2) (2024) 3940-3951, http://dx.doi.org/10.1109/TIV.2024.3357056..

[11]

W. Zhou, et al., CACFNet: Cross-modal attention cascaded fusion network for RGB-T urban scene parsing, IEEE Trans. Intell. Veh. 9 (1) (2023) 1919-1929.

[12]

J. Huang, et al., DepthMatch: Semi-supervised RGB-D scene parsing through depth-guided regularization, IEEE Signal Process. Lett. 32 (2025) 2549-2553, https://doi.org/10.1109/LSP.2025.3575640.

[13]

K. He, et al., Deep Residual Learning for Image Recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2016, pp. 770-778.

[14]

Z. Liu, et al., Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, 2021, pp. 10012-10022.

[15]

G. Tang, et al., TiCoSS: Tightening the coupling between semantic segmentation and stereo matching within a joint learning framework, IEEE Trans. Autom. Sci. Eng. 22 (2025) 18646-18658, https://doi.org/10.1109/TASE.2025.3586286.

[16]

Z. Liu, et al., A ConvNet for the 2020s, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2022, pp. 11976-11986.

[17]

M.-J. Lee, et al., SG-RoadSeg+: End-to-end freespace detection upgraded at data, feature, and loss levels, IEEE Trans. Instrum. Meas. 74 (2025) 1-9, https://doi.org/10.1109/TIM.2025.3579733.

[18]

Y. Feng, et al., SNE-RoadSegV2: Advancing heterogeneous feature fusion and fallibility awareness for freespace detection, IEEE Trans. Instrum. Meas. 74 (2025) 1-9.

[19]

H. Bao, et al., BEiT: BERT Pre-Training of Image Transformers, in: International Conference on Learning Representations, ICLR, 2022.

[20]

Z. Peng, et al., BEiT v2: Masked image modeling with vector-quantized visual tokenizers, 2022, arXiv preprint arXiv:2208.06366.

[21]

M. Oquab, et al., DINOv2: Learning robust visual features without supervision, Trans. Mach. Learn. Res., (2023).

[22]

A. Krizhevsky, et al., ImageNet classification with deep convolutional neural networks, Adv. Neural Inf. Process. Syst. (NeurIPS) 25 (2012) 1097-1105.

[23]

S. Zhao, et al., Mitigating modality discrepancies for RGB-T semantic segmentation, IEEE Trans. Neural Networks Learn. Syst., (2023), 10.1109/TNNLS.2022.3233089.

[24]

F. Deng, et al., FEANet: Feature-enhanced attention network for RGB-Thermal real-time semantic segmentation, 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS, IEEE, 2021, pp. 4467-4473.

[25]

J. Liu, et al., Revisiting modality-specific feature compensation for visible-infrared person re-identification, IEEE Trans. Circuits Syst. Video Technol. 32 (10) (2022) 7226-7240.

[26]

D. Seichter, et al., Efficient RGB-D semantic segmentation for indoor scene analysis, 2021 IEEE International Conference on Robotics and Automation, ICRA, IEEE, 2021, pp. 13525-13531.

[27]

Q. Zhang, et al., RGB-T salient object detection via fusing multi-level CNN features, IEEE Trans. Image Process. 29 (2019) 3321-3335.

[28]

S. Y, et al., FuseSeg: Semantic segmentation of urban scenes based on RGB and thermal data fusion, IEEE Trans. Autom. Sci. Eng. 18 (3) (2020) 1000-1011.

[29]

J. Huang, et al., RoadFormer+: Delivering RGB-X scene parsing through scale-aware information decoupling and advanced heterogeneous feature fusion, IEEE Trans. Intell. Veh. 10 (5) (2025) 3156-3165.

[30]

X. Zhu, et al., Deformable DETR: Deformable Transformers for End-to-End Object Detection, in: International Conference on Learning Representations, ICLR, 2020.

[31]

S. Sun, et al., ReMaX: Relaxing for better training on efficient panoptic segmentation, Adv. Neural Inf. Process. Syst. (NeurIPS), 36 (2024).

[32]

Y.-H. Kim, et al., MS-UDA: Multi-spectral unsupervised domain adaptation for thermal image semantic segmentation, IEEE Robot. Autom. Lett. 6 (4) (2021) 6497-6504.

[33]

N. Silberman, et al., Indoor segmentation and support inference from RGBD images, European Conference on Computer Vision, ECCV, Springer, 2012, pp. 746-760.

[34]

J. Long, et al., Fully Convolutional Networks for Semantic Segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2015, pp. 3431-3440.

[35]

L.-C. Chen, et al., DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs, IEEE Trans. Pattern Anal. Mach. Intell. 40 (4) (2017) 834-848.

[36]

L. Chen, et al., Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation, in: Proceedings of the European Conference on Computer Vision, ECCV, 2018, pp. 801-818.

[37]

T.-Y. Lin, et al., Feature Pyramid Networks for Object Detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2017, pp. 2117-2125.

[38]

A. Vaswani, et al., Attention is all you need, Adv. Neural Inf. Process. Syst. (NeurIPS) 30 (2017) 5998-6008.

[39]

S. Zheng, et al., Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2021, pp. 6881-6890.

[40]

E. Xie, et al., SegFormer: Simple and efficient design for semantic segmentation with transformers, Adv. Neural Inf. Process. Syst. (NeurIPS) 34 (2021) 12077-12090.

[41]

J. Dai, et al., Convolutional Feature Masking for Joint Object and Stuff Segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2015, pp. 3992-4000.

[42]

B. Hariharan, et al., Simultaneous detection and segmentation, Proceedings of the European Conference on Computer Vision, ECCV, Springer, 2014, pp. 297-312.

[43]

B. Cheng, et al., Per-pixel classification is not all you need for semantic segmentation, Adv. Neural Inf. Process. Syst. (NeurIPS) 34 (2021) 17864-17875.

[44]

B. Cheng, et al., Masked-attention Mask Transformer for Universal Image Segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2022, pp. 1290-1299.

[45]

W. Zhou, et al., DBCNet: Dynamic bilateral cross-fusion network for RGB-T urban scene understanding in intelligent vehicles, IEEE Trans. Syst. Man, Cybern.: Syst. 53 (12) (2023) 7631-7641.

[46]

G. Huang, et al., Densely Connected Convolutional Networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2017, pp. 4700-4708.

[47]

J. Zhang, et al., CMX: Cross-modal fusion for RGB-X semantic segmentation with transformers, IEEE Trans. Intell. Transp. Syst. 24 (12) (2023) 14679-14694.

[48]

Z. J, et al., Delivering Arbitrary-Modal Semantic Segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2023, pp. 1136-1147.

[49]

K. He, et al., Masked Autoencoders Are Scalable Vision Learners, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2022, pp. 16000-16009.

[50]

T. Xiao, et al., Unified Perceptual Parsing for Scene Understanding, in: Proceedings of the European Conference on Computer Vision, ECCV, 2018, pp. 418-434.

[51]

B. Yin, et al., DFormer: Rethinking RGBD representation learning for semantic segmentation, Int. Conf. Learn. Represent. (ICLR), (2024).

[52]

K. Li, et al., UniFormer: Unifying convolution and self-attention for visual recognition, IEEE Trans. Pattern Anal. Mach. Intell. 45 (10) (2023) 12581-12600.

[53]

S. Xie, et al., Aggregated Residual Transformations for Deep Neural Networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2017, pp. 1492-1500.

[54]

J. Li, et al., RoadFormer: Duplex transformer for RGB-normal semantic road scene parsing, IEEE Trans. Intell. Veh. 9 (7) (2024) 5163-5172.

[55]

R. Bachmann, et al., MultiMAE: Multi-modal multi-task masked autoencoders, European Conference on Computer Vision, ECCV, Springer, 2022, pp. 348-367.

[56]

S. Huang, et al., FaPN: Feature-Aligned Pyramid Network for Dense Image Prediction, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, ICCV, 2021, pp. 864-873.

[57]

Z. Chen, et al., Vision Transformer Adapter for Dense Predictions, in: The Eleventh International Conference on Learning Representations, ICLR, 2023.

[58]

X. Yang, et al., PolyMaX: General Dense Prediction with Mask Transformer, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, WACV, 2024, pp. 1050-1061.

[59]

N. Carion, et al., End-to-end object detection with transformers, European Conference on Computer Vision, ECCV, Springer, 2020, pp. 213-229.

[60]

M. Cordts, et al., The Cityscapes Dataset for Semantic Urban Scene Understanding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2016, pp. 3213-3223.

[61]

I. Loshchilov, F. Hutter, Decoupled Weight Decay Regularization, in: International Conference on Learning Representations, ICLR, 2018.

[62]

W. Zhou, et al., Edge-Aware Guidance Fusion Network for RGB–Thermal Scene Parsing, Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, (3) 2022, pp. 3571–3579.

[63]

W. Zhou, et al., Embedded control gate fusion and attention residual learning for RGB–thermal urban scene parsing, IEEE Trans. Intell. Transp. Syst. 24 (5) (2023) 4794-4803.

[64]

X. He, et al., SFAF-MA: Spatial feature aggregation and fusion with modality adaptation for RGB-Thermal semantic segmentation, IEEE Trans. Instrum. Meas. 72 (2023) 1-10.

[65]

X. Guo, et al., Low-light enhancement and global-local feature interaction for RGB-T semantic segmentation, IEEE Trans. Instrum. Meas., (2025).

[66]

R. Girdhar, et al., OMNIVORE: A Single Model for Many Visual Modalities, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2022, pp. 16102-16112.

[67]

S. Srivastava, G. Sharma, OmniVec: Learning Robust Representations With Cross Modal Sharing, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, WACV, 2024, pp. 1236-1248.

[68]

Y. Wang, et al., Multimodal Token Fusion for Vision Transformers, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2022, pp. 12186-12195.

[69]

S. Du, et al., AsymFormer: Asymmetrical cross-modal representation learning for mobile platform real-time RGB-D semantic segmentation, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2024, pp. 7608-7615.

[70]

X. Chen, et al., Bi-directional cross-modality feature propagation with separation-and-aggregation gate for RGB-D semantic segmentation, European Conference on Computer Vision, ECCV, Springer, 2020, pp. 561-577.

[71]

H. Touvron, et al., Training data-efficient image transformers & distillation through attention, International Conference on Machine Learning, ICML, PMLR, 2021, pp. 10347-10357.

[72]

A.P. Steiner, et al., How to train your ViT? Data, augmentation, and regularization in vision transformers, Trans. Mach. Learn. Res., (2022).

[73]

A. Dosovitskiy, et al., An image is worth 16x16 words: Transformers for image recognition at scale, Int. Conf. Learn. Represent. (ICLR), (2020).

PDF (3064KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/