A Deep Learning System for Automatic Localization of Anatomical Landmarks in X-rays to Assist in Diagnosis and Surgical Planning

Hui Zhang , Tengfei Li , Ahmad Alenezi , Xinguo Wang , Qinsheng Hu , Zekun Jiang , Quan Wei

BIO Integration ›› 2026, Vol. 7 ›› Issue (1) : 34

PDF (3715KB)
BIO Integration ›› 2026, Vol. 7 ›› Issue (1) :34 DOI: 10.15212/bioi-2026-0075
Original Article
research-article
A Deep Learning System for Automatic Localization of Anatomical Landmarks in X-rays to Assist in Diagnosis and Surgical Planning
Author information +
History +
PDF (3715KB)

Abstract

Background: Accurate localization of anatomical landmarks is crucial for clinical diagnosis and treatment assessment. However, existing convolutional neural network (CNN)-based methods may result in global spatial information loss and consequent localization failures in the presence of complex anatomical structures or parenchymal abnormalities. Therefore, a method capable of modeling global context while preserving local information is needed.

Methods: Leveraging the Transformer’s ability to capture long-range dependencies, we propose a novel landmark localization framework, Res-SwinFusion, which integrates a Swin Transformer and a classical CNN backbone in parallel. To effectively merge their complementary features, we designed a feature interactive aggregation module that fuses semantic representations from both branches. Additionally, we introduced a discrimination feature guidance module to provide pixel-level cues and disambiguate landmark locations. We further analyzed the effects of various Gaussian heatmap settings on convergence.

Results: Res-SwinFusion achieved strong performance across three anatomical landmark localization datasets. The mean radial errors were 1.04 mm and 1.37 mm on two public cephalogram test sets, 0.63 mm on a public hand X-ray dataset, and 1.44 mm on an internal pelvic X-ray dataset. Ablation studies indicated that Transformer-based global modeling, feature interactive aggregation, and discrimination feature guidance each contributed to improved localization accuracy.

Conclusion: The proposed Res-SwinFusion framework offers a solution for anatomical landmark localization with enhanced robustness and precision by combining global contextual modeling and local feature preservation. Code is publicly available at https://github.com/JZK00/Res-SwinFusion.

Keywords

Anatomical landmarks / automatic localization / deep learning / X-ray

Cite this article

Download citation ▾
Hui Zhang, Tengfei Li, Ahmad Alenezi, Xinguo Wang, Qinsheng Hu, Zekun Jiang, Quan Wei. A Deep Learning System for Automatic Localization of Anatomical Landmarks in X-rays to Assist in Diagnosis and Surgical Planning. BIO Integration, 2026, 7 (1) : 34 DOI:10.15212/bioi-2026-0075

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Liu Y, Song Y, Du Y, Qin C, Xu T. From language to world: bridging LLMs and world models for intelligent surgery. Innov Inform. 2026; 2(2): 100040. [DOI: 10.59717/j.xinn-inform.2026.100040]

[2]

Alam F, Rahman SU, Ullah S, Gulati K. Medical image registration in image guided surgery: issues, challenges and research opportunities. Biocybern Biomed Eng. 2018; 38(1): 71-89. [DOI: 10.1016/j.bbe.2017.10.001]

[3]

Bier B, Goldmann F, Zaech JN, Fotouhi J, Hegeman R, et al. Learning to detect anatomical landmarks of the pelvis in X-rays from arbitrary views. Int J Comput Assist Radiol Surg. 2019; 14(9): 1463-73. [PMID: 31006106 DOI: 10.1007/s11548-019-01975-5]

[4]

Štern D, Payer C, Lepetit V, Urschler M. Automated age estimation from hand MRI volumes using deep learning. In: Ourselin S, Joskowicz L, Sabuncu M, Unal G, Wells W, editors. Medical Image Computing and Computer-Assisted Intervention - MICCAI 2016. Cham: Springer; 2016. pp. 194-202. [DOI: 10.1007/978-3-319-46723-8_23]

[5]

Zheng Y, John M, Liao R, Boese J, Kirschstein U, et al. Automatic aorta segmentation and valve landmark detection in C-arm CT: application to aortic valve implantation. In: Jiang T, Navab N, Pluim JPW, Viergever MA, editors. Medical Image Computing and Computer-Assisted Intervention - MICCAI 2010. Berlin, Heidelberg: Springer; 2010. pp. 476-83. [DOI: 10.1007/978-3-642-15705-9_58]

[6]

Huang Y, Fan F, Syben C, Roser P, Mill L, et al. Cephalogram synthesis and landmark detection in dental cone-beam CT systems. Med Image Anal. 2021; 70: 102028. [PMID: 33744833 DOI: 10.1016/j.media.2021.102028]

[7]

Ibragimov B, Likar B, Pernuš F, Vrtovec T. Shape representation for efficient landmark-based segmentation in 3-D. IEEE Trans Med Imaging. 2014; 33(4): 861-74. [PMID: 24710155 DOI: 10.1109/TMI.2013.2296976]

[8]

Kamoen A, Dermaut L, Verbeeck R. The clinical significance of error measurement in the interpretation of treatment results. Eur J Orthod. 2001; 23(5): 569-78. [PMID: 11668876 DOI: 10.1093/ejo/23.5.569]

[9]

Lindner C, Wang CW, Huang CT, Li CH, Chang SW, et al. Fully automatic system for accurate localisation and analysis of cephalometric landmarks in lateral cephalograms. Sci Rep. 2016; 6: 33581. [PMID: 27645567 DOI: 10.1038/srep33581]

[10]

Grau V, Alcañiz M, Juan MC, Monserrat C, Knoll C. Automatic localization of cephalometric landmarks. J Biomed Inform. 2001; 34(3): 146-56. [PMID: 11723697 DOI: 10.1006/jbin.2001.1014]

[11]

El-Feghi I, Sid-Ahmed MA, Ahmadi M. Automatic localization of craniofacial landmarks for assisted cephalometry. Pattern Recognit. 2004; 37(3): 609-21. [DOI: 10.1016/j.patcog.2003.09.002]

[12]

Mirzaalian H, Hamarneh G. Automatic globally-optimal pictorial structures with random decision forest based likelihoods for cephalometric X-ray landmark detection. In Automatic Cephalometric X-ray Landmark Detection Challenge 2014, in Conjunction with IEEE International Symposium on Biomedical Imaging; 2014. pp. 1-12.

[13]

Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, et al. A survey on deep learning in medical image analysis. Med Image Anal. 2017; 42: 60-88. [PMID: 28778026 DOI: 10.1016/j.media.2017.07.005]

[14]

Arık , Ibragimov B, Xing L. Fully automated quantitative cephalometry using convolutional neural networks. J Med Imaging (Bellingham). 2017; 4(1): 014501. [PMID: 28097213 DOI: 10.1117/1.JMI.4.1.014501]

[15]

Zhong Z, Li J, Zhang Z, Jiao Z, Gao X. An attention-guided deep regression model for landmark detection in cephalograms. In: Shen D, et al., editors. Medical Image Computing and Computer Assisted Intervention. Cham: Springer; 2019. pp. 540-8. [DOI: 10.1007/978-3-030-32226-7_60]

[16]

He T, Yao J, Tian W, Yi Z, Tang W, et al. Cephalometric landmark detection by considering translational invariance in the two-stage framework. Neurocomput. 2021; 464: 15-26. [DOI: 10.1016/j.neucom.2021.08.042]

[17]

Zeng M, Yan Z, Liu S, Zhou Y, Qiu L. Cascaded convolutional networks for automatic cephalometric landmark detection. Med Image Anal. 2021; 68: 101904. [PMID: 33290934 DOI: 10.1016/j.media.2020.101904]

[18]

He X, Zhou Y, Zhao J, Zhang D, Yao R, et al. Swin transformer embedding UNet for remote sensing image semantic segmentation. IEEE Transact Geosci Remote Sens. 2022; 60: 1-15. [DOI: 10.1109/TGRS.2022.3144165]

[19]

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, et al. Attention is all you need. In: von Luxburg U, Guyon I, Bengio S, Wallach H, Fergus R, editors. Advances in neural information processing systems, Vol. 30. Red Hook, NY: Curran Associates, Inc.; 2017. pp. 3104-14.

[20]

Li J, Wang W, Chen C, Zhang T, Zha S, et al. TransBTSV2: Towards better and more efficient volumetric segmentation of medical images. arXiv preprint arXiv:2201.12785. 2022.

[21]

Chen J, Lu Y, Yu Q, Luo X, Zhou Y, et al. TransUNet: transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306. 2021.

[22]

Cao H, Wang Y, Chen J, Jiang D, Zhang X, et al. Swin-Unet: Unet-like pure transformer for medical image segmentation. In: Karlinsky L, Michaeli T, Nishino K, editors. ECCV 2022: Proceedings of the Computer Vision - ECCV 2022 Workshops, Part III; 2022 Oct 23-27; Tel Aviv, Israel. Cham: Springer; 2023. pp, 205-18. [DOI: 10.1007/978-3-031-25066-8_9]

[23]

Lin A, Chen B, Xu J, Zhang Z, Lu G. DS-TransUNet: dual swin transformer U-Net for medical image segmentation. IEEE Trans Instrum Meas. 2022; 71: 1-15. [DOI: 10.1109/TIM.2022.3178991]

[24]

Zhou HY, Guo J, Zhang Y, Yu L, Wang L, et al. nnFormer: interleaved transformer for volumetric segmentation. arXiv preprint arXiv:2109.03201. 2021.

[25]

Zhang Y, Liu H, Hu Q. TransFuse: fusing transformers and CNNs for medical image segmentation. In: de Bruijne M, Cattin PC, Cotin S, Padoy N, Speidel S, et al., editors. MICCAI 2021: Medical Image Computing and Computer Assisted Intervention. Cham: Springer; 2021. pp. 14-24. [DOI: 10.1007/978-3-030-87193-2_2]

[26]

Hong W, Kim SM, Choi J, Paeng JY, Mun JH, et al. Deep reinforcement learning using a multi-scale agent with a normalized reward strategy for automatic cephalometric landmark detection. In: 2023 4th International Conference on Big Data Analytics and Practices (IBDAP). Bangkok, Thailand: IEEE; 2023. pp. 1-6. [DOI: 10.1109/IBDAP58581.2023.10271989]

[27]

Zhu Z, Gu X, Dong L, Liu Y, Wang Y, et al. An ensemble-based deep learning method through multi-scale cross-attention training for cephalometric landmark localization on lateral X-ray images. 2025. [DOI: 10.21203/rs.3.rs-6105085/v1]

[28]

Lu G, Zhang Y, Kong Y, Zhang C, Coatrieux JL, et al. Landmark localization for cephalometric analysis using multiscale image patch-based graph convolutional networks. IEEE J Biomed Health Inform. 2022; 26(7): 3015-24. [PMID: 35259123 DOI: 10.1109/JBHI.2022.3157722]

[29]

Lee H, Park M, Kim J. Cephalometric landmark detection in dental x-ray images using convolutional neural networks. Proc. SPIE 10134, Medical Imaging 2017: Computer-Aided Diagnosis, 101341W; 2017. pp. 494-9. [DOI: 10.1117/12.2255870]

[30]

Noothout JMH, De Vos BD, Wolterink JM, Postma EM, Smeets PAM, et al. Deep learning-based regression and classification for automatic landmark localization in medical images. IEEE Trans Med Imaging. 2020; 39(12): 4011-22. [PMID: 32746142 DOI: 10.1109/TMI.2020.3009002]

[31]

Payer C, Štern D, Bischof H, Urschler M. Regressing heatmaps for multiple landmark localization using CNNs. In: Ourselin S, Joskowicz L, Sabuncu M, Unal G, Wells W, editors. Medical Image Computing and Computer-Assisted Intervention - MICCAI 2016. Cham: Springer; 2016. pp. 230-8. [DOI: 10.1007/978-3-319-46723-8_27]

[32]

Payer C, Štern D, Bischof H, Urschler M. Integrating spatial configuration into heatmap regression based CNNs for landmark localization. Med Image Anal. 2019; 54: 207-19. [PMID: 30947144 DOI: 10.1016/j.media.2019.03.007]

[33]

Thaler F, Payer C, Urschler M, Štern D. Modeling annotation uncertainty with gaussian heatmaps in landmark localization. J Mach Learn Biomed Imaging. 2021; 14: 1-27. [DOI: 10.59275/j.melba.2021-77a7]

[34]

Chen R, Ma Y, Chen N, Lee D, Wang W. Cephalometric landmark detection by attentive feature pyramid fusion and regression-voting. In: Shen D, et al., editors. International Conference on Medical Image Computing and Computer-Assisted Intervention - MICCAI 2019. Cham: Springer; 2019. pp. 873-81. [DOI: 10.1007/978-3-030-32248-9_97]

[35]

Shamshad F, Khan S, Zamir SW, Khan MH, Hayat M, et al. Transformers in medical imaging: a survey. Med Image Anal. 2023; 88: 102802. [DOI: 10.1016/j.media.2023.102802]

[36]

Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, et al. An image is worth 16x16 words: transformers for image recognition at scale. International Conference on Learning Representations. arXiv:2010.11929v2. 2021. [DOI: 10.48550/arXiv.2010.11929]

[37]

Liu Z, Lin Y, Cao Y, Hu H, Wei Y, et al. Swin transformer: hierarchical vision transformer using shifted windows. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada: IEEE; 2021. pp. 9992-10002. [DOI: 10.1109/ICCV48922.2021.00986]

[38]

Yang S, Quan Z, Nie M, Yang W. TransPose: keypoint localization via transformer. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada: IEEE; 2021. pp. 11802-12. [DOI: 10.1109/ICCV48922.2021.01159]

[39]

Li Y, Zhang S, Wang Z, Yang S, Zhou E. TokenPose: learning keypoint tokens for human pose estimation. In: 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada: IEEE; 2021, pp. 11313-22. [DOI: 10.1109/ICCV48922.2021.01112]

[40]

Li H, Guo Z, Rhee SM, Han S, Han JJ. Towards accurate facial landmark detection via cascaded transformers. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, LA, USA: IEEE; 2022. pp. 4176-85. [DOI: 10.1109/CVPR52688.2022.00414]

[41]

Zhu H, Yao Q, Zhou SK. DATR: domain-adaptive transformer for multi-domain landmark detection. arXiv:2203.06433v1. 2022. [DOI: 10.48550/arXiv.2203.06433]

[42]

Ao Y, Hong W. Swin transformer combined with convolutional encoder for cephalometric landmarks detection. In: 2021 18th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP). Chengdu, China: IEEE; 2021. pp. 184-7. [DOI: 10.1109/ICCWAMTIP53232.2021.9674147]

[43]

Shaker A, Maaz M, Rasheed H, Khan S, Yang MH, et al. UNETR++: delving into efficient and accurate 3D medical image segmentation. IEEE Trans Med Imaging. 2024; 43(9): 3377-90. [PMID: 38722726 DOI: 10.1109/TMI.2024.3398728]

[44]

Jin Z, Qiu Y, Zhang K, Li H, Luo W. MB-TaylorFormer V2: improved multi-branch linear transformer expanded by Taylor formula for image restoration. IEEE Trans Pattern Anal Mach Intell. 2025; 47(7): 5990-6005. [PMID: 40208767 DOI: 10.1109/TPAMI.2025.3559891]

[45]

Zhang K, Li D, Luo W, Ren W, Liu W. Enhanced spatio-temporal interaction learning for video deraining: faster and better. IEEE Trans Pattern Anal Mach Intell. 2023; 45(1): 1287-93. [PMID: 35130145 DOI: 10.1109/TPAMI.2022.3148707]

[46]

Zhang K, Li R, Yu Y, Luo W, Li C. Deep dense multi-scale network for snow removal using semantic and depth priors. IEEE Trans Image Process. 2021; 30: 7419-7431. [DOI: 10.1109/TIP.2021.3104166]

[47]

Zhang K, Luo W, Zhong Y, Ma L, Liu W, et al. Adversarial spatio-temporal learning for video deblurring. IEEE Trans Image Process. 2019; 28: 291-301. [DOI: 10.1109/TIP.2018.2867733]

[48]

Lin A, Chen B, Xu J, Zhang Z, Lu G, et al. DS-transUNet: dual swin transformer U-net for medical image segmentation. IEEE Trans Instrum Meas. 2022; 71: 1-15. [DOI: 10.1109/TIM.2022.3178991]

[49]

Ronneberger O, Fischer P, Brox T. U-Net: convolutional networks for biomedical image segmentation. In: Navab N, Hornegger J, Wells W, Frangi A, editors. International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer; 2015. pp. 234-41. [DOI: 10.1007/978-3-319-24574-4_28]

[50]

Hu J, Shen L, Sun G, Albanie S, Wu E. Squeeze-and-excitation networks. Proc. IEEE Conf Comput Vis Pattern Recognit; 2018. pp. 7132-41. [DOI: 10.48550/arXiv.1709.01507]

[51]

Woo S, Park J, Lee JY, Kweon IS. CBAM: convolutional block attention module. In: Ferrari V, Hebert M, Sminchisescu C, Weiss Y, editors. Computer Vision - ECCV 2018. Cham: Springer; 2018. pp. 3-19. [DOI: 10.1007/978-3-030-01234-2_1]

[52]

Wang CW, Huang CT, Lee JH, Li CH, Chang SW, et al. A benchmark for comparison of dental radiography analysis algorithms. Med Image Anal. 2016; 31: 63-76. [PMID: 26974042 DOI: 10.1016/j.media.2016.02.004]

[53]

Zhu H, Yao Q, Xiao L, Zhou SK. You only learn once: universal anatomical landmark detection. In: de Bruijne M, et al. Medical Image Computing and Computer Assisted Intervention. Cham: Springer; 2021. pp. 85-95. [DOI: 10.1007/978-3-030-87240-3_9]

[54]

Zhou J, Wang Y, Huang C, Dai C, Tan C. CeLR: a transformer-based regression network for accurate cephalometric landmark detection in high-resolution X-ray imaging. IEEE Trans Med Imaging. 2026; 45(5): 2283-94. [DOI: 10.1109/TMI.2026.3652170]

[55]

Laitenberger F, Scheuer HT, Scheuer HA, Lilienthal E, You S, et al. Cephalometric landmark detection using vision transformers with direct coordinate prediction. J Craniomaxillofac Surg. 2025; 53(9): 1518-29. [PMID: 40603150 DOI: 10.1016/j.jcms.2025.05.021]

[56]

Polizzi A, Leonardi R. Automatic cephalometric landmark identification with artificial intelligence: an umbrella review of systematic reviews. J Dent. 2024; 146: 105056. [PMID: 38729291 DOI: 10.1016/j.jdent.2024.105056]

[57]

Joham SJ, Hadzic A, Urschler M. Implicit is not enough: explicitly enforcing anatomical priors inside landmark localization models. Bioengineering (Basel). 2024; 11(9): 932. [PMID: 39329674 DOI: 10.3390/bioengineering11090932]

[58]

Song C, Jeong Y, Huh H, Park JW, Paeng JY, et al. Multi-scale 3D cephalometric landmark detection based on direct regression with 3D CNN architectures. Diagnostics (Basel). 2024; 14(22): 2605. [PMID: 39594271 DOI: 10.3390/diagnostics14222605]

[59]

Gillot M, Miranda F, Baquero B, Ruellas A, Gurgel M, et al. Automatic landmark identification in cone-beam computed tomography. Orthod Craniofac Res. 2023; 26(4): 560-7. [PMID: 36811276 DOI: 10.1111/ocr.12642]

[60]

Baldini B, Rubiu G, Serafin M, Bologna M, Facchi GM, et al. Automated 3D cephalometry: a lightweight V-net for landmark localization on CBCT. Comput Med Imaging Graph. 2026; 128: 102700. [PMID: 41519031 DOI: 10.1016/j.compmedimag.2026.102700]

[61]

Deitermann M, Pankert T, Jaganathan S, Röhrle O, Hölzle F, et al. Automated detection of mandibular landmarks in CT data using a dual-input approach in a two-stage design. Comput Methods Programs Biomed. 2026; 273: 109113. [PMID: 41086721 DOI: 10.1016/j.cmpb.2025.109113]

[62]

Dhruba DD, Goetz S, Pria OFD, Reith T, Reutzel A, et al. Deep learning-based cardiac MRI planning from localizers to cine views using landmark detection. Acad Radiol. 2026; 33(3): 924-35. [PMID: 41353071 DOI: 10.1016/j.acra.2025.11.028]

[63]

Tao R, Ye K, Zhang W, Sun W, Yu D, et al. X2P-Net: context-aware 2D/3D vertebra localization. Bioengineering (Basel). 2026; 13(2): 178. [PMID: 41749718 DOI: 10.3390/bioengineering13020178]

[64]

Ma J, He Y, Li F, Han L, You C, et al. Segment anything in medical images. Nat Commun. 2024; 15(1): 654. [PMID: 38253604 DOI: 10.1038/s41467-024-44824-z]

[65]

Urschler M, Ebner T, Štern D. Integrating geometric configuration and appearance information into a unified framework for anatomical landmark localization. Med Image Anal. 2018; 43: 23-36. [PMID: 28963961 DOI: 10.1016/j.media.2017.09.003]

PDF (3715KB)

4

Accesses

0

Citation

Detail

Sections
Recommended

/