A comparative study of deep learning-based retinal image registration methods

Thenuka Dharmaseelan , Neelabh Sinha , Samyyia Ashraf , Kimia Daneshvar , Amit John , Periklis Giannakis , Yik Ting Chan , Yiu Wai Chan , Nikolas Pontikos

Exploration of Digital Health Technologies ›› 2026, Vol. 4 ›› Issue (1) : 101194

PDF (6902KB)
Exploration of Digital Health Technologies ›› 2026, Vol. 4 ›› Issue (1) :101194 DOI: 10.37349/edht.2026.101194
Original Article
research-article
A comparative study of deep learning-based retinal image registration methods
Author information +
History +
PDF (6902KB)

Abstract

Aim: To benchmark three deep learning-based retinal image registration methods RetinaRegNet, EyeLiner, and GeoFormer on the Fundus Image Registration (FIRE) dataset to compare registration accuracy and computational efficiency using mean landmark error (MLE) as the primary outcome measure.

Methods: The three image registration approaches were evaluated using the FIRE dataset under consistent conditions across varying image overlap conditions (Classes S, A, and P). These included: (a) RetinaRegNet, which incorporates diffusion features, dual keypoint sampling through Scale-Invariant Feature Transform (SIFT) and random, two-stage outlier removal, and a multilevel registration hierarchy progressing from homography to polynomial transforms; (b) EyeLiner, which integrates anatomical segmentation with SuperPoint feature extraction, LightGlue matching, and thin-plate spline warping; (c) GeoFormer, which builds on Local Feature Transformers (LoFTR) through cross-attention mechanisms and Random Sampling Consensus (RANSAC)-based refinement. Registration performance was quantified using MLE.

Results: Across all 134 FIRE image pairs, RetinaRegNet achieved the lowest overall MLE (3.12 pixels), outperforming EyeLiner (3.81 pixels) and GeoFormer (6.06 pixels). Class-specific analysis showed that RetinaRegNet delivered the highest accuracy in Class S images (1.70 pixels), competitive performance in Class A (5.24 pixels), and the strongest results in the most challenging Class P cases (4.57 pixels). GeoFormer demonstrated the shortest processing time at 0.32 seconds per image pair, compared with 4.92 seconds for EyeLiner and 31.23 seconds for RetinaRegNet. In Class P, RetinaRegNet achieved a 59.2% improvement in accuracy relative to GeoFormer (4.57 vs 11.20 pixels). The code is available at: https://github.com/ThenukaDharmaseelan/image_Registration.

Conclusions: Overall, the evaluation reveals a clear trade-off between registration precision and computational speed. RetinaRegNet achieves the lowest MLE for complex clinical cases despite higher computational cost. EyeLiner balances precision and speed for routine use, while GeoFormer prioritizes rapid throughput where processing speed is critical.

Keywords

retinal image registration / fundus photography / deep learning / EyeLiner / GeoFormer / RetinaRegNet

Cite this article

Download citation ▾
Thenuka Dharmaseelan, Neelabh Sinha, Samyyia Ashraf, Kimia Daneshvar, Amit John, Periklis Giannakis, Yik Ting Chan, Yiu Wai Chan, Nikolas Pontikos. A comparative study of deep learning-based retinal image registration methods. Exploration of Digital Health Technologies, 2026, 4 (1) : 101194 DOI:10.37349/edht.2026.101194

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Ramsey DJ, Sunness JS, Malviya P, Applegate C, Hager GD, Handa JT. Automated image alignment and segmentation to follow progression of geographic atrophy in age-related macular degeneration. Retina. 2014; 34:1296-307.

[2]

Hussain MA, Govindaiah A, Souied E, Smith RT, Bhuiyan A. Automated tracking and change detection for age-related macular degeneration progression using retinal fundus imaging. 2018 Joint 7th International Conference on Informatics, Electronics & Vision (ICIEV) and 2018 2nd International Conference on Imaging, Vision & Pattern Recognition (icIVPR); 2018 Jun 25-29; Kitakyushu, Japan. IEEE; pp. 394-8.

[3]

Balakrishnan G, Zhao A, Sabuncu MR, Guttag J, Dalca AV. VoxelMorph: A Learning Framework for Deformable Medical Image Registration. IEEE Trans Med Imaging. 2019.

[4]

Sokooti H, de Vos B, Berendsen F, Lelieveldt BPF, Išgum I, Staring M. Nonrigid Image Registration Using Multi-scale 3D Convolutional Neural Networks. In: Descoteaux M, Maier-Hein L, Franz A, Jannin P, Collins D, Duchesne S, editors. Medical Image Computing and Computer Assisted Intervention − MICCAI 2017. MICCAI 2017. Cham: Springer; 2017.

[5]

Zhang J, Wang Y, Dai J, Cavichini M, Bartsch DG, Freeman WR, et al. Two-Step Registration on Multi-Modal Retinal Images via Deep Neural Networks. IEEE Trans Image Process. 2022; 31:823-38.

[6]

Zhang J, Wang Y, Bartsch DG, Freeman WR, Nguyen TQ, An C. Perspective Distortion Correction for Multi-Modal Registration between Ultra-Widefield and Narrow-Angle Retinal Images. Annu Int Conf IEEE Eng Med Biol Soc. 2021; 2021:4086-91.

[7]

Zhang J, An C, Dai J, Amador M, Bartsch DU, Borooah S. Joint Vessel Segmentation and Deformable Registration on Multi-Modal Retinal Images Based on Style Transfer. 2019 IEEE International Conference on Image Processing (ICIP); 2019 Sep 22-25; Taipei, Taiwan. IEEE; 2019. pp. 839-43.

[8]

Benvenuto GA, Colnago M, Casaca W. Unsupervised Deep Learning Network for Deformable Fundus Image Registration. ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2022 May 23-27; Singapore, Singapore. IEEE; 2022. pp. 1281-5.

[9]

Tian Y, Hu Y, Ma Y, Hao H, Mou L, Yang J, et al. Multi-scale U-net with Edge Guidance for Multimodal Retinal Image Deformable Registration. Annu Int Conf IEEE Eng Med Biol Soc. 2020; 2020:1360-3.

[10]

Zou B, He Z, Zhao R, Zhu C, Liao W, Li S. Non-rigid retinal image registration using an unsupervised structure-driven regression network. Neurocomputing. 2020; 404:14-25.

[11]

Liu J, Li X, Wei Q, Xu J, Ding D. Semi-Supervised Keypoint Detector and Descriptor for Retinal Image Matching. In: Avidan S, Brostow G, Cissé M, Farinella GM, Hassner T, editors. Computer Vision - ECCV 2022. ECCV 2022. Lecture Notes in Computer Science. 2022. pp. 593-609.

[12]

Wang AQ, Yu EM, Dalca AV, Sabuncu MR. A robust and interpretable deep learning framework for multi-modal registration via keypoints. Med Image Anal. 2023; 90:102962.

[13]

Chen J, Frey EC, He Y, Segars WP, Li Y, Du Y. TransMorph: Transformer for unsupervised medical image registration. Med Image Anal. 2022; 82:102615.

[14]

Sun J, Shen Z, Wang Y, Bao H, Zhou X. LoFTR: Detector-Free Local Feature Matching with Transformers. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition; 2021 Jun 20-25; Nashville, USA. IEEE; 2021. pp. 8918-27.

[15]

Ho J, Jain A, Abbeel P. Denoising Diffusion Probabilistic Models. Adv Neural Inf Process Syst. 2020.

[16]

Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-Resolution Image Synthesis with Latent Diffusion Models. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. 2022 Jun 18-24; New Orleans, USA. IEEE; 2022. pp. 10674-85.

[17]

Sivaraman VB, Imran M, Wei Q, Muralidharan P, Tamplin MR, Grumbach IM, et al. RetinaRegNet: A zero-shot approach for retinal image registration. Comput Biol Med. 2025; 186:109645.

[18]

Veturi YA, McNamara S, Kinder S, Clark CW, Thakuria U, Bearce B, et al. EyeLiner: A Deep Learning Pipeline for Longitudinal Image Registration Using Fundus Landmarks. Ophthalmol Sci. 2024; 5:100664.

[19]

Liu J, Li X. Geometrized Transformer for Self-Supervised Homography Estimation. Proceedings of the IEEE International Conference on Computer Vision. 2023 October 1-6; Paris, France. IEEE; 2023. pp. 9522-31.

[20]

Hernandez-Matas C, Zabulis X, Triantafyllou A, Anyfanti P, Douma S, Argyros AA. FIRE: Fundus Image Registration dataset. Model Artif Intell Ophthalmol. 2017; 1:16-28.

[21]

Zhou Y, Wagner SK, Chia MA, Zhao A, Woodward-Court P, Xu M, et al. AutoMorph: Automated Retinal Vascular Morphology Quantification Via a Deep Learning Pipeline. Transl Vis Sci Technol. 2022; 11:12.

[22]

Cheng B, Schwing AG, Kirillov A. Per-Pixel Classification is Not All You Need for Semantic Segmentation. Adv Neural Inf Process Syst. 2021; 22:17864-75.

[23]

Detone D, Malisiewicz T, Rabinovich A. SuperPoint: Self-Supervised Interest Point Detection and Description. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2018 Jun 18-22; Salt Lake City, USA. IEEE; 2018. pp. 337-49.

[24]

Lindenberger P, Sarlin PE, Pollefeys M. LightGlue: Local Feature Matching at Light Speed. 2023 IEEE/CVF International Conference on Computer Vision (ICCV); 2023 Oct 1-6; Paris, France. IEEE; 2023. pp. 17581-92.

[25]

Lin TY, Dollar P, Girshick R, He K, Hariharan B, Belongie S. Feature Pyramid Networks for Object Detection. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21-26; Honolulu, HI, USA. IEEE; 2017. pp. 936-44.

PDF (6902KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉