FineSegNeRF: Semantic NeRF for fine semantic segmentation

Weitao Liu , Shiyao Wang , Junjun Wu , Qinghua Lu

Biomimetic Intelligence and Robotics ›› 2026, Vol. 6 ›› Issue (3) : 100302

PDF (2324KB)
Biomimetic Intelligence and Robotics ›› 2026, Vol. 6 ›› Issue (3) :100302 DOI: 10.1016/j.birob.2026.100302
Research Article
research-article
FineSegNeRF: Semantic NeRF for fine semantic segmentation
Author information +
History +
PDF (2324KB)

Abstract

Semantic segmentation is a crucial technology for intelligent vehicles, enabling robust scene understanding in complex driving environments. However, existing methods often struggle with small, distant, and overlapping objects, posing challenges for safe autonomous operation. To address these limitations, we present FineSegNeRF, a model designed for fine semantic segmentation of such challenging scenarios. Our approach separates features along the depth dimension to perceive stereoscopic scene from spatial dimension, and then uses NeRF’s multi-view consistency to optimize the separated features for fine understanding. Meanwhile, a new “Semantic Uncertainty Neural Volume Render” method is proposed for constraining the consistency of volume density and semantic uncertainty estimation to further improve the semantic segmentation performance. Compared to current representative RGB-D and NeRF fusion semantic segmentation methods, our approach performs remarkable competitiveness in terms of fine semantic segmentation on VKITTI 2 and Replica datasets.

Keywords

Fine semantic segmentation / NeRF / Multi-view fusion / RGB-D semantic segmentation / Stereo perception

Cite this article

Download citation ▾
Weitao Liu, Shiyao Wang, Junjun Wu, Qinghua Lu. FineSegNeRF: Semantic NeRF for fine semantic segmentation. Biomimetic Intelligence and Robotics, 2026, 6 (3) : 100302 DOI:10.1016/j.birob.2026.100302

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Fabian Brickwedde, Steffen Abraham, Rudolf Mester, Mono-sf: Multi-view geometry meets single-view depth for monocular scene flow estimation of dynamic traffic scenes, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2780-2790.

[2]

Yan Xu, Xinge Zhu, Jianping Shi, Guofeng Zhang, Hujun Bao, Hongsheng Li, Depth completion from sparse lidar data with depth-normal constraints, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2811-2820.

[3]

Lucas Teixeira, Martin R. Oswald, Marc Pollefeys, Margarita Chli, Aerial single-view depth completion with image-guided uncertainty estimation, IEEE Robot. Autom. Lett. 5 (2) (2020) 1055-1062.

[4]

Nicolas Marchal, Charlotte Moraldo, Hermann Blum, Roland Siegwart, Cesar Cadena, Abel Gawel, Learning densities in feature space for reliable segmentation of indoor scenes, IEEE Robot. Autom. Lett. 5 (2) (2020) 1032-1038.

[5]

Caner Hazirbas, Lingni Ma, Csaba Domokos, Daniel Cremers, Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture, Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part I 13, Springer, 2017, pp. 213-228.

[6]

Yanhua Cheng, Rui Cai, Zhiwei Li, Xin Zhao, Kaiqi Huang, Locality-sensitive deconvolution networks with gated fusion for RGB-D indoor semantic segmentation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 3029-3037.

[7]

Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang, Wayne Wu, Chen Qian, Hongsheng Li, Gang Zeng, Bi-directional cross-modality feature propagation with separation-and-aggregation gate for RGB-D semantic segmentation, European Conference on Computer Vision, Springer, 2020, pp. 561-577.

[8]

Xinxin Hu, Kailun Yang, Lei Fei, Kaiwei Wang, Acnet: Attention based network to exploit complementary features for rgbd semantic segmentation, 2019 IEEE International Conference on Image Processing, ICIP, IEEE, 2019, pp. 1440-1444.

[9]

Yikai Wang, Fuchun Sun, Ming Lu, Anbang Yao, Learning deep multimodal feature representation with asymmetric multi-layer fusion, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 3902-3910.

[10]

David Eigen, Rob Fergus, Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture, in: Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2650-2658.

[11]

Shu Kong, Charless C. Fowlkes, Recurrent scene parsing with perspective understanding in the loop, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 956-965.

[12]

Vladimir Nekrasov, Thanuja Dharmasiri, Andrew Spek, Tom Drummond, Chunhua Shen, Ian Reid, Real-time joint semantic segmentation and depth estimation using asymmetric annotations, 2019 International Conference on Robotics and Automation, ICRA, IEEE, 2019, pp. 7101-7107.

[13]

Ling Zhou, Zhen Cui, Chunyan Xu, Zhenyu Zhang, Chaoqun Wang, Tong Zhang, Jian Yang, Pattern-structure diffusion for multi-task learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 4514-4523.

[14]

Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng, Nerf: Representing scenes as neural radiance fields for view synthesis, Commun. ACM 65 (1) (2021) 99-106.

[15]

Ricardo Martin-Brualla, Noha Radwan, Mehdi S.M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, Daniel Duckworth, Nerf in the wild: Neural radiance fields for unconstrained photo collections, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7210-7219.

[16]

Alex Yu, Vickie Ye, Matthew Tancik, Angjoo Kanazawa, pixelnerf: Neural radiance fields from one or few images, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4578-4587.

[17]

Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P. Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, Thomas Funkhouser, Ibrnet: Learning multi-view image-based rendering, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4690-4699.

[18]

Alex Trevithick, Bo Yang, Grf: Learning a general radiance field for 3d representation and rendering, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 15182-15192.

[19]

Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, Hao Su, Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14124-14133.

[20]

Ajay Jain, Matthew Tancik, Pieter Abbeel, Putting nerf on a diet: Semantically consistent few-shot view synthesis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 5885-5894.

[21]

Khawaja Iftekhar Rashid, Chenhui Yang, Chenxi Huang, Fast-DSAGCN: Enhancing semantic segmentation with multifaceted attention mechanisms, Neurocomputing 587 (2024) 127625.

[22]

Wenbin Zou, Guoguang Hua, Yue Zhuang, Shishun Tian, Real-Time Passable Area segmentation with consumer RGB-D cameras for the visually impaired, IEEE Trans. Instrum. Meas. 72 (2023) 1-11.

[23]

Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, Andrew J. Davison, In-place scene labelling and understanding with implicit scene representation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 15838-15847.

[24]

Andreas Geiger, Philip Lenz, Christoph Stiller, Raquel Urtasun, Vision meets robotics: The kitti dataset, Int. J. Robot. Res. 32 (11) (2013) 1231-1237.

[25]

Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, Oscar Beijbom, nuscenes: A multimodal dataset for autonomous driving, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11621-11631.

[26]

Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, Matthias Nießner, Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5828-5839.

[27]

Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al., Habitat: A platform for embodied ai research, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 9339-9347.

[28]

Xinyu Huang, Peng Wang, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, Ruigang Yang, The apolloscape open dataset for autonomous driving and its application, IEEE Trans. Pattern Anal. Mach. Intell. 42 (10) (2019) 2702-2719.

[29]

Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, Bernt Schiele, The cityscapes dataset for semantic urban scene understanding, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3213-3223.

[30]

Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, Trevor Darrell, BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2020, pp. 2633-2642.

[31]

Nathan Silberman, Derek Hoiem, Pushmeet Kohli, Rob Fergus, Indoor segmentation and support inference from rgbd images, Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12, Springer, 2012, pp. 746-760.

[32]

Yohann Cabon, Naila Murray, Martin Humenberger, Virtual kitti 2, 2020, arXiv preprint arXiv:2001.10773.

[33]

Olaf Ronneberger, Philipp Fischer, Thomas Brox, U-net: Convolutional networks for biomedical image segmentation, Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, Springer, 2015, pp. 234-241.

[34]

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778.

[35]

Rasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, Henrik Aanæs, Large scale multi-view stereopsis evaluation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 406-413.

[36]

Ilya Loshchilov, Frank Hutter, SGDR: Stochastic Gradient Descent with Warm Restarts, in: International Conference on Learning Representations, 2017, Published in the ICLR 2017 Conference Proceedings.

[37]

Fangfu Liu, Chubin Zhang, Yu Zheng, Yueqi Duan, Semantic Ray: Learning a Generalizable Semantic Field with Cross-Reprojection Attention, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17386-17396.

[38]

Saurabh Gupta, Ross Girshick, Pablo Arbeláez, Jitendra Malik, Learning rich features from RGB-D images for object detection and segmentation, Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, Springer, 2014, pp. 345-360.

[39]

Bowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu, Ming-Ming Cheng, Qibin Hou, DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation, in: The Twelfth International Conference on Learning Representations, 2024.

[40]

Jonathan Long, Evan Shelhamer, Trevor Darrell, Fully convolutional networks for semantic segmentation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 3431-3440.

[41]

Xiaojuan Qi, Renjie Liao, Jiaya Jia, Sanja Fidler, Raquel Urtasun, 3d graph neural networks for rgbd semantic segmentation, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 5199-5208.

[42]

Yang Zhang, Chenyun Xiong, Junjie Liu, Xuhui Ye, Guodong Sun, Spatial-information guided adaptive context-aware network for efficient RGB-D semantic segmentation, IEEE Sensors J., (2023).

[43]

Yikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang, Fuchun Sun, Yunhe Wang, Multimodal token fusion for vision transformers, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12186-12195.

[44]

Haoming Liu, Li Guo, Zhongwen Zhou, Hanyuan Zhang, Pyramid-context guided feature fusion for RGB-D semantic segmentation, 2022 IEEE International Conference on Multimedia and Expo Workshops, ICMEW, IEEE, 2022, pp. 1-6.

[45]

Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis, et al., 3D gaussian splatting for real-time radiance field rendering., ACM Trans. Graph., 42 (4), (2023),139–1.

[46]

Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, Jifeng Dai, Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers, European Conference on Computer Vision, Springer, 2022, pp. 1-18.

PDF (2324KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/