Currently, foundation model had attracted a lot attention for monocular depth estimation in endoscopic surgery. However, degradation in clinical scenarios is often complex for endoscopic image, which leads to compromised robustness. Therefore, we propose a self-supervised generative feature driven EndoMDNet that aims to mitigate degradation for robust monocular depth estimation in endoscopic surgery. Specifically, a content recognition mechanism is designed to guide diffusion model to generate detail information that is utilized as the supplementation of degraded depth feature. Moreover, diffusion model tends to generate artifacts that may also be inconsistent with the target distribution. To address this problem, we propose wavelet-based refined adaptive fusion block (WRAF) to filter noise of generative feature and adaptively fuse it with degraded depth feature. Finally, extensive experiment on screen for child anxiety related emotional disorders (SCARED), stereoscopic endoscopic reconstruction validation-CT (SERV-CT) and Hamlyn datasets demonstrate the robustness of our proposed method.
| [1] |
Zha R, Cheng X, Li H, et al. . Endosurf: neural surface reconstruction of deformable tissues with stereo endoscope videos. Intemational Conference on Medical Image Computing and Computer-Assisted Intervention, October 8–12, 2023, Vancouver, BC, Canada. 2023, Cham, Springer Nature Switzerland1323[C]
|
| [2] |
Yang L, Kang B, Huang Z, et al. . Depth anything: unleashing the power of large-scale unlabeled data. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 17–21, 2024, Seattle, WA, USA. 2024, New York, IEEE1037110381[C]
|
| [3] |
Shao S, Pei Z, Chen W, et al. . Self-supervised monocular depth and EGO-motion estimation in endoscopy: appearance flow to the rescue. Medical image analysis, 2022, 77: 102338 J]
|
| [4] |
Cui B, Islam M, Bai L, et al. . EndoDAC: efficient adapting foundation model for self-supervised depth estimation from any endoscopic camera. International Conference on Medical Image Computing and Computer-Assisted Intervention, October 6–10, 2024, Marrakesh, Morocco. 2024, Cham, Springer Nature Switzerland208-218[C]
|
| [5] |
Lei Y J, Wang Y, Chan S X, et al. . BEDiff: denoising diffusion probabilistic models for building extraction. Optoelectronics letters, 2025, 21: 298-305 J]
|
| [6] |
Li G, Rao C, Mo J, et al. . Rethinking diffusion model for multi-contrast MRI super-resolution. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 17–21, 2024, Seattle WA, USA. 2024, New York, IEEE1136511374[C]
|
| [7] |
Zhang N, Nex F, Vosselman G, et al. . Lite-mono: a lightweight CNN and transformer architecture for self-supervised monocular depth estimation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 18–22, 2023, Vancouver, Canada. 2023, New York, IEEE1853718546[C]
|
| [8] |
Ulyanov D, Vedaldi A, Lempitsky V. Improved texture networks: maximizing quality and diversity in feed-forward stylization and texture synthesis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, July 21–16, 2017, Hawaii, USA. 2017, New York, IEEE69246932[C]
|
| [9] |
Ioffe S, Szegedy C. Batch normalization: accelerating deep network training by reducing internal covariate shift. International Conference on Machine Learning, July 6–11, 2015, Lille, France. 2015448456[C]
|
| [10] |
Kong L, Xie S, Hu H, et al. . Robodepth: robust out-of-distribution depth estimation under corruptions. Advances in neural information processing systems, 2023, 36: 21298-21342 J]
|
| [11] |
Jaspers T J, Boers T G, Kusters C H, et al. . Robustness evaluation of deep neural networks for endoscopic image analysis: insights and strategies. Medical image analysis, 2024, 94: 103157 J]
|
| [12] |
ALLAN M, MCLEOD J, WANG C, et al. Stereo correspondence and reconstruction of endoscopic data challenge: version 4[EB/OL]. (2021-01-28) [2025-04-10]. https://arxiv.org/abs/2101.01133.
|
| [13] |
Edwards P, Psychogyios D, Speidel S, et al. . SERV-CT: a disparity dataset from CT for validation of endoscopic 3D reconstruction. Medical image analysis, 2021, 76102302 J]
|
| [14] |
Recasens D, Lamarca J, Fácil J M, et al. . Endo-depth-and-motion: reconstruction and tracking in endoscopic videos using depth networks and photometric constraints. IEEE robotics and automation letters, 2021, 6(4): 7225-7232 J]
|
| [15] |
Godard C, Mac Aodha O, Firman M, et al. . Digging into self-supervised monocular depth estimation. Proceedings of the IEEE/CVF International Conference on Computer Vision, June 16–20, 2019, Long Beach, CA, USA. 2019, New York, IEEE3828-3838[C]
|
| [16] |
Ozyoruk K B, Gokceler G I, Bobrow T L, et al. . EndoSLAM dataset and an unsupervised monocular visual odometry and depth estimation approach for endoscopic videos. Medical image analysis, 2021, 71102058 J]
|
| [17] |
Bian J W, Zhan H, Wang N, et al. . Unsupervised scale-consistent depth learning from video. International journal of computer vision, 2021, 129(9): 2548-2564 J]
|
RIGHTS & PERMISSIONS
Tianjin University of Technology