A Review of Artificial Intelligence in Ophthalmology: Key Aspects, Challenges, and Future Directions

Partha Pratim Ray

Eye & ENT Research ›› 2026, Vol. 3 ›› Issue (3) : 125 -154.

PDF (1190KB)
Eye & ENT Research ›› 2026, Vol. 3 ›› Issue (3) :125 -154. DOI: 10.1002/eer3.70049
REVIEW ARTICLE
A Review of Artificial Intelligence in Ophthalmology: Key Aspects, Challenges, and Future Directions
Author information +
History +
PDF (1190KB)

Abstract

Artificial intelligence (AI) is increasingly reshaping ophthalmology because the specialty depends heavily on structured imaging, quantitative measurements, and repeatable diagnostic workflows. This review provides a clinically grounded and translationally oriented synthesis of AI in ophthalmology, covering methodological foundations, ophthalmic imaging modalities, public datasets, disease‐specific applications, evaluation metrics, deployment barriers, and future directions. Unlike reviews that mainly summarize algorithmic performance by disease category or model type, this article organizes ophthalmic AI through an integrated framework that emphasizes clinical use cases, evidence maturity, translational readiness, and real‐world implementation requirements. The review examines applications across population screening, referral triage, disease grading, progression monitoring, prognosis, treatment guidance, workflow support, and automated reporting. Major disease domains include diabetic retinopathy, glaucoma, age‐related macular degeneration, cataract, infectious keratitis, and keratoconus. Particular attention is given to the distinction between retrospective proof‐of‐concept studies, external validation, multicenter evaluation, prospective trials, and real‐world deployment. The review also interprets evaluation metrics from a clinical perspective, highlighting the importance of threshold selection, sensitivity, specificity, false referral burden, missed disease, calibration, uncertainty, segmentation adequacy, robustness, and generalization. Key translational challenges include dataset bias, domain shift, interpretability, privacy, regulatory oversight, infrastructure constraints, workflow integration, and post‐deployment monitoring. Emerging paradigms such as multimodal AI, foundation models, generative AI, and edge‐based point‐of‐care systems are discussed cautiously, with emphasis on hallucination risk, clinical grounding, accountability, and the gap between benchmark performance and deployment readiness. Overall, the review argues that the next phase of ophthalmic AI should move beyond high accuracy values toward prospective validation, external generalization, clinician‐centered design, calibrated uncertainty, and accountable integration into real‐world eye‐care pathways.

Graphical abstract

Keywords

artificial intelligence / clinical decision support / deep learning / diabetic retinopathy / explainable AI / glaucoma / ophthalmology / optical coherence tomography / retinal imaging

Cite this article

Download citation ▾
Partha Pratim Ray. A Review of Artificial Intelligence in Ophthalmology: Key Aspects, Challenges, and Future Directions. Eye & ENT Research, 2026, 3 (3) : 125-154 DOI:10.1002/eer3.70049

登录浏览全文

4963

注册一个新账户 忘记密码

1 Introduction

Ophthalmology is one of the most fertile clinical domains for artificial intelligence (AI), owing to its strong reliance on imaging, standardized diagnostic pathways, and quantitatively interpretable anatomical structures [1]. Many ophthalmic diseases manifest through visible or measurable changes in the retina, optic nerve head, macula, cornea, lens, or ocular vasculature. As a result, ophthalmic practice generates a large volume of structured visual and functional data that can be systematically analyzed using computational methods. Modalities such as color fundus photography, optical coherence tomography (OCT), OCT angiography (OCTA), slit‐lamp imaging, corneal topography or tomography, and visual field (VF) testing provide high‐resolution information that is particularly suitable for automated pattern recognition, segmentation, grading, and longitudinal monitoring [2]. In contrast to several other medical specialties where clinically relevant information is often dispersed across heterogeneous records, free‐text notes, laboratory results, and subjective assessments, ophthalmology frequently offers a more direct relationship between digital data and disease characterization. This makes the field especially suitable for the development, validation, and deployment of AI‐based diagnostic and decision‐support systems [3].

The clinical need for such systems is substantial. Vision‐threatening conditions, including diabetic retinopathy (DR), glaucoma, age‐related macular degeneration (AMD), cataract, keratoconus, and infectious or inflammatory corneal diseases, contribute significantly to the global burden of visual impairment and preventable blindness [4]. Many of these disorders are either asymptomatic in the early stages or require repeated specialist assessment for timely detection and monitoring. Early diagnosis and appropriate intervention are therefore essential to prevent irreversible visual loss. However, access to ophthalmic expertise remains uneven, particularly in rural, low‐resource, and underserved settings. The increasing prevalence of diabetes, population aging, shortage of trained eye‐care professionals, and growing demand for screening and follow‐up services further intensify pressure on ophthalmology systems. In this context, AI has emerged as a potentially scalable approach to support population screening, referral triage, disease grading, diagnosis, progression monitoring, treatment planning, and workflow optimization [5].

Technological progress in ophthalmic AI has been rapid and transformative. Early approaches based on handcrafted features and classical machine learning (ML) established important foundations by using explicit descriptors such as vessel caliber, lesion morphology, optic disc parameters, texture patterns, and corneal curvature indices. Although these methods offered interpretability and clinical alignment, they were limited by the need for manual feature engineering, task‐specific design, and reduced generalizability across devices and populations. The emergence of deep learning (DL), particularly convolutional neural networks (CNNs), shifted the field toward automated representation learning directly from raw or minimally processed images. This led to substantial improvements in classification, detection, segmentation, and grading tasks across DR, glaucoma, AMD, cataract, infectious keratitis, and keratoconus. More recently, attention mechanisms, vision transformers, self‐supervised learning, multimodal learning, and foundation models have expanded the scope of ophthalmic AI beyond isolated image‐based prediction toward broader forms of clinical intelligence, including automated reporting, risk prediction, longitudinal assessment, and decision support.

Despite these advances, high algorithmic performance in retrospective datasets does not automatically translate into clinical value. Many ophthalmic AI studies continue to rely on internally validated datasets, selected image‐quality conditions, limited demographic diversity, and benchmark‐oriented evaluation. Such studies are useful for technical development but may not fully capture the variability of real‐world clinical practice. Important translational questions remain unresolved: whether models generalize across imaging devices and populations, whether predictions remain calibrated under domain shift, whether uncertainty is communicated appropriately, whether outputs can be safely integrated into clinical workflows, and whether AI‐assisted care improves patient outcomes rather than merely increasing diagnostic throughput. These concerns are particularly important in ophthalmology because the consequences of false reassurance, delayed referral, excessive false positives, or poorly explained recommendations may directly affect visual prognosis.

Unlike earlier reviews that mainly summarized disease‐specific AI performance, this review critically maps ophthalmic AI across clinical use cases, evidence maturity, translational readiness, and deployment bottlenecks. It integrates methodological foundations, disease‐specific evidence, imaging modalities, public datasets, evaluation metrics, real‐world validation, and emerging paradigms such as foundation models, generative AI, multimodal systems, and edge deployment within a clinically prioritized framework. In doing so, the review emphasizes not only what ophthalmic AI systems can achieve under experimental conditions, but also what evidence is required before they can be responsibly implemented in routine eye care.

Accordingly, this review is organized around both technical and clinical perspectives. First, it introduces the major methodological foundations of ophthalmic AI, including machine learning, DL, CNNs, vision transformers, multimodal models, foundation models, generative AI, and explainable AI. Second, it examines the principal ophthalmic imaging modalities and representative public datasets that have shaped the development of AI systems. Third, it reviews disease‐specific applications in DR, glaucoma, AMD, cataract, infectious keratitis, and keratoconus, while highlighting differences in evidence maturity across these domains. Fourth, it discusses evaluation metrics not merely as mathematical indicators, but as clinically meaningful tools for understanding screening safety, referral burden, calibration, robustness, and segmentation adequacy. Finally, it addresses translational challenges, scenario‐based deployment requirements, regulatory and ethical considerations, and future directions for clinically accountable ophthalmic AI. Figure 1 presents the overall architecture followed in this review. It links ophthalmic data sources, preprocessing and governance, AI model development, computational functions, clinical use cases, disease‐specific domains, validation hierarchy, translational readiness, deployment pathways, and post‐deployment monitoring.

2 Math Review Methodology, Literature Search Strategy, and Scope

This review was designed as a structured narrative synthesis of AI in ophthalmology, with emphasis on methodological foundations, imaging modalities, disease‐specific applications, evaluation practices, translational barriers, and future clinical directions. The purpose was not only to summarize reported model performance, but also to critically examine the maturity of evidence, the clinical relevance of different AI applications, and the major requirements for safe deployment in ophthalmic practice.

Relevant literature was identified through searches of major scientific databases, including PubMed/MEDLINE, Scopus, Web of Science, IEEE Xplore, ScienceDirect, SpringerLink, and Google Scholar. The search focused primarily on peer‐reviewed articles published during the recent period of rapid development in ophthalmic AI, while also including older landmark studies and foundational papers where necessary. Representative search terms included combinations of: “artificial intelligence,” “machine learning,” “deep learning,” “ophthalmology,” “retinal imaging,” “fundus photography,” “optical coherence tomography,” “OCT angiography,” “diabetic retinopathy,” “glaucoma,” “age‐related macular degeneration,” “cataract,” “keratitis,” “keratoconus,” “explainable AI,” “foundation models,” “multimodal AI,” “clinical validation,” and “AI deployment.” Studies were considered relevant if they addressed AI‐based analysis, screening, diagnosis, grading, segmentation, progression assessment, treatment support, workflow integration, or translational evaluation in ophthalmology. Priority was given to studies reporting clinically meaningful tasks, clearly defined datasets, validation strategies, diagnostic or segmentation performance, external testing, prospective evaluation, or real‐world deployment. Public dataset papers, benchmark studies, major disease‐specific AI studies, regulatory and ethical discussions, and high‐impact review articles were also considered. Studies with limited ophthalmic relevance, unclear methodology, duplicate reporting, purely promotional claims, or insufficient description of model evaluation were not emphasized.

The included literature was organized across complementary dimensions. First, studies were grouped according to methodological approach, including classical ML, DL, CNNs, vision transformers, multimodal models, foundation models, generative AI, and explainable AI. Second, they were classified according to ophthalmic data modality, including fundus photography, OCT, OCTA, slit‐lamp and anterior segment imaging, corneal topography or tomography, and VF testing. Third, disease‐specific evidence was synthesized for DR, glaucoma, AMD, cataract, infectious keratitis, and keratoconus. Finally, studies were interpreted in relation to clinical use cases, including screening, referral triage, disease grading, longitudinal monitoring, prognosis, treatment guidance, and workflow support.

Because ophthalmic AI is a rapidly evolving field, this review emphasizes not only algorithmic accuracy but also evidence quality, clinical validity, robustness, calibration, workflow compatibility, and deployment readiness. Reported performance values were interpreted cautiously, especially when derived from retrospective single‐center datasets or internally validated experiments. Greater weight was assigned to studies involving external validation, multicenter data, prospective testing, real‐world evaluation, or clinically integrated workflows. This approach was adopted to distinguish technical feasibility from translational readiness and to provide a clinically meaningful assessment of the current state of AI in ophthalmology.

2.1 Evidence Quality, Validation Hierarchy, and Risk of Bias

The evidentiary strength of ophthalmic AI studies varies considerably, and this variation must be considered when interpreting reported performance. High accuracy in a controlled experimental setting does not necessarily imply clinical readiness. Therefore, AI systems in ophthalmology should be assessed according to a validation hierarchy that distinguishes early proof‐of‐concept development from externally validated, prospectively tested, and clinically deployed systems. Retrospective single‐center studies generally represent the lowest level of translational evidence. Such studies are useful for initial algorithm development and feasibility assessment, but they are often limited by homogeneous patient populations, restricted imaging devices, local acquisition protocols, and dataset‐specific characteristics. Models trained and tested within the same institutional environment may learn patterns that are specific to that setting rather than generalizable disease features. As a result, performance reported in retrospective single‐center studies may overestimate real‐world reliability.

External validation is an essential next step because it evaluates model performance on data collected from a different institution, population, device, or clinical environment. A model that maintains performance under external validation provides stronger evidence of generalizability than one evaluated only on internal test data. However, external validation alone is not sufficient for full clinical translation. External datasets may still be retrospectively curated, may not reflect routine clinical workflows, and may exclude poor‐quality images, atypical presentations, or operational constraints commonly encountered in practice. Multicenter validation further improves the assessment of generalizability by exposing the model to broader variability in patient demographics, imaging devices, disease prevalence, acquisition protocols, and clinical labeling practices. Such validation is particularly important in ophthalmology because image characteristics may differ substantially across camera manufacturers, OCT platforms, screening programs, geographic regions, and levels of clinical expertise. Multicenter testing can therefore reveal performance instability that may remain hidden in single‐center or narrowly selected datasets.

Prospective clinical trials provide a stronger level of evidence because they evaluate AI systems under real or near‐real clinical conditions. Unlike retrospective studies, prospective evaluations can assess not only diagnostic performance but also workflow impact, clinician interaction, referral behavior, time efficiency, patient acceptability, and safety outcomes. In ophthalmology, prospective studies are especially important for screening systems, triage tools, and progression‐monitoring models, where the clinical value of AI depends on how the output affects actual decision‐making and follow‐up. Real‐world deployment requires continued monitoring even after initial validation. AI systems may experience performance drift over time due to changes in imaging devices, software updates, patient population, disease prevalence, referral pathways, or clinical practice patterns. Post‐market or post‐deployment monitoring is therefore necessary to detect degradation, subgroup‐specific errors, calibration failure, and unintended workflow consequences. Such monitoring should include periodic auditing, error analysis, clinician feedback, and mechanisms for safe model updating where appropriate.

Risk of bias is another major concern in ophthalmic AI. Bias may arise from device type, ethnicity, geography, age distribution, disease severity, image quality, disease prevalence, referral patterns, and annotation practices. For example, a model trained mainly on high‐quality images from one camera type may perform poorly on images from portable devices used in community screening. Similarly, models trained on populations from high‐resource settings may not generalize to rural, underserved, or ethnically diverse populations. Referral bias may also occur when datasets are drawn from tertiary centers, where disease severity is higher than in general screening populations.

Bias can also be introduced through labeling procedures. Ground‐truth labels may vary according to the expertise of graders, diagnostic criteria, imaging modality, or availability of confirmatory tests. In diseases such as glaucoma, infectious keratitis, and keratoconus, diagnostic labels may be less straightforward than in standardized DR grading. Inter‐observer variability, uncertain cases, and incomplete clinical information can therefore influence both training and evaluation.

For these reasons, ophthalmic AI studies should report dataset composition, imaging devices, acquisition settings, disease prevalence, demographic characteristics, labeling procedures, exclusion criteria, and subgroup performance wherever possible. Transparent reporting allows readers to judge whether a model is likely to generalize beyond the study setting. It also helps distinguish between algorithmic performance under idealized conditions and clinical usefulness under real‐world variability. Figure 2 presents the pathway distinguishes early retrospective proof‐of‐concept studies from externally validated, multicenter, prospectively tested, clinically deployed, and continuously monitored AI systems.

2.2 Related Works

Recent literature on AI in ophthalmology has expanded rapidly, covering multimodal learning, foundation models, agentic AI, glaucoma diagnosis, corneal nerve imaging, clinician and patient perceptions, optometric education, and precision medicine. This section summarizes representative recent works and positions the present review in relation to these studies.

Chen et al. [6] reviewed the transition from visual question answering (VQA) to intelligent AI agents in ophthalmology. Their work emphasizes that ophthalmic practice requires the integration of heterogeneous clinical data and interactive decision‐making, which creates limitations for traditional task‐specific AI systems. VQA combines computer vision and natural language processing (NLP) to interpret ophthalmic images through user‐driven questions, while multimodal AI agents extend this capability through continuous dialog, tool use, and context‐aware clinical decision support. The review highlights the role of large language models (LLMs) in improving reasoning, adaptability, and task execution. However, it also identifies important limitations, including limited multimodal datasets, lack of standardized evaluation protocols, and difficulties in clinical integration. The authors argue that realizing the full potential of LLM‐driven ophthalmic AI will require close collaboration between AI researchers and the ophthalmic community.

Vamsidhar et al. [7] presented a systematic review of multimodal AI in ophthalmology, focusing on methods, applications, and future directions. Their review covers developments from 2020 to 2025 and emphasizes the integration of structural, functional, and clinical data, including fundus photography, OCT, OCTA, and clinical metadata. The authors show that multimodal AI can improve detection, grading, and prognostication in DR, glaucoma, and macular degeneration. They discuss major fusion strategies, including early fusion, late fusion, hybrid fusion, and transformer‐based fusion models, and relate these methods to precision ophthalmology. The review also identifies persistent challenges, including lack of standardized datasets, imaging‐protocol variability, interpretability gaps, privacy concerns, regulatory barriers, and limited real‐world implementation. It further discusses federated learning, explainable AI, and generative AI as possible approaches to address data scarcity and improve trust in AI‐assisted diagnostics.

Siwik et al. [8] reviewed the application of artificial intelligence in glaucoma diagnosis. They highlight glaucoma as a chronic, progressive, and potentially blinding disease in which early‐stage detection is especially important because symptoms may be unnoticed by patients. Traditional glaucoma diagnosis relies on multiple examinations, including intraocular pressure measurement, gonioscopy, fundus examination, OCT, and VF testing. The authors note that these procedures can be time‐consuming and may involve subjective interpretation. Their review analyzes literature from the past decade and discusses the fundamentals of AI, the effectiveness of algorithms in glaucoma diagnosis, and the assessment of disease risk. The study concludes that AI may support earlier detection of individuals at risk, reduce vision loss, reduce physician burden, and improve quality of care, while also acknowledging limitations in diagnostic reliability and clinical applicability.

Soetikno et al. [9] examined the emergence of agentic AI and its potential implications for ophthalmic research and clinical practice. Their review explains that advances in LLMs have enabled AI systems capable of autonomously performing complex scientific tasks, including peer review, hypothesis generation, systematic reviews, experimental design, and biomedical analysis. Although direct ophthalmology‐specific applications remain at an early stage, the authors argue that ophthalmology is well positioned for agentic AI adoption because of its data‐rich nature. Potential applications include automated chart review, health economics modeling, and enhanced image analysis. The article presents agentic AI as a paradigm shift in scientific research, with potential to improve productivity, rigor, and innovation. At the same time, it stresses the need for ethical governance, authorship clarity, privacy protection, bias mitigation, accountability, rigorous validation, and interdisciplinary training.

Barcelo‐Canton et al. [10] reviewed applications of AI in corneal nerve image analysis. Their article focuses on in vivo confocal microscopy (IVCM), which enables high‐resolution visualization of corneal nerves at a microscopic level. Traditionally, corneal nerve images require manual examination, which is time‐consuming and labor‐intensive. The authors describe how AI enables automatic and semiautomatic identification, segmentation, and quantification of corneal nerve parameters. They also discuss the diagnostic relevance of AI‐assisted corneal nerve analysis in dry eye disease and neuropathic corneal pain. The review notes that AI improves reproducibility, reduces operator dependency, and shortens analysis time. It also highlights the use of AI for identifying microneuromas and detecting changes in corneal nerve metrics. However, the authors emphasize that large‐scale validation is needed before widespread implementation and that future multimodal AI integration may further improve diagnostic accuracy and disease stratification.

Garcia et al. [11] investigated perceptions of AI in retina clinical care among retina specialists and patients through a multicenter survey. The study included 291 patients and 78 physicians across five retina practices. Their findings showed important differences between physician and patient attitudes toward AI, especially regarding patient autonomy, desire for AI use, and responsibility for harm caused by AI. Patients were more likely than physicians to believe they should be able to choose whether AI is used in their care, while physicians reported greater desire for AI use in clinical care. Both groups expressed discomfort with AI taking direct clinical roles, such as deciding treatments or answering questions about diagnosis and treatment. The authors conclude that as AI enters retina care, developers and clinicians must understand and address patient concerns, particularly regarding autonomy, responsibility, and the limits of AI involvement in clinical decision‐making.

Buckmaster et al. [12] studied attitudes and knowledge levels of optometry students and educators toward AI in optometric practice using an online cross‐sectional survey. The survey included 254 respondents from 14 of 15 optometry universities in the United Kingdom, including 213 students and 41 educators. The study found that most students had not received AI training, and many reported little or no knowledge of AI applications in optometry. Educators were more likely than students to have received AI training and reported higher knowledge levels. Respondents who had undergone AI training showed more positive attitudes toward AI in optometry. Educators strongly supported including AI in the optometry curriculum. The study identifies a need for future AI educational initiatives in optometry programs, while also noting that further research is needed to guide curriculum development and training priorities.

Gharbi et al. [13] reviewed foundation models for ophthalmic imaging. Their survey systematically examined 12 ophthalmic foundation models developed between 2022 and July 2025. The authors analyze the evolution of these models in terms of modality integration, including unimodal, multimodal, and vision‐language approaches; pretraining objectives, including generative and contrastive strategies; and supervision methods, including image‐guided and text‐guided learning. They report a clear shift from domain‐specific unimodal models toward modality‐agnostic foundation models guided by clinical text. The review also discusses emerging techniques such as imaging‐modality‐agnostic encoders, synthetic data augmentation, and computationally efficient architectures. The authors conclude that future work should focus on broader modality integration, higher‐dimensional spatial and temporal inputs, diverse pretraining strategies, and standardized benchmark datasets.

Pratap et al. [14] reviewed DL technology in genomics, radiotherapy, and ophthalmology within the broader framework of precision medicine. Their review presents precision medicine as a shift from generalized treatment toward individualized care informed by genetic, environmental, and lifestyle factors. The authors discuss how AI and DL enable integration of heterogeneous biomedical data. In genomics, AI supports variant classification, gene‐expression modeling, and multiomics fusion for disease‐risk prediction. In radiology and biomedical imaging, convolutional and transformer‐based architectures enhance lesion detection, image reconstruction, and radiogenomic mapping. The review also recognizes emerging applications in ophthalmic imaging and digital pathology for early diagnosis and personalized therapy planning. It identifies major challenges, including data interoperability, transparency, ethical governance, and trustworthy implementation, and discusses federated learning, explainable AI, and privacy‐preserving computation as possible solutions.

Tang et al. [15] reviewed recent advances and future directions of AI in glaucoma management. Their work describes glaucoma as a leading cause of irreversible blindness and emphasizes the diagnostic and management challenges created by its insidious progression and irreversible late‐stage optic nerve damage. The review discusses how ML, DL, and LLMs are transforming glaucoma care across screening, precise diagnosis, treatment optimization, and long‐term patient management. The authors highlight multidimensional innovations, including image analysis, molecular biomarker identification, treatment‐response prediction, and surgical planning. They also discuss key challenges and future developments, emphasizing that AI in glaucoma is moving beyond diagnosis toward broader management support.

Grzybowski et al. [16] discussed agentic AI in ophthalmology as a movement toward autonomous, adaptive, and ethical eye care. Their perspective defines agentic AI as a new generation of systems capable of autonomous goal‐directed reasoning, dynamic decision‐making, and coordinated action. Unlike conventional task‐specific algorithms, agentic AI systems can perceive, plan, and act within complex healthcare environments with limited supervision. The authors argue that ophthalmology is an ideal field for this paradigm because it is image‐intensive and data‐rich. They outline possible future applications, including intelligent diagnostic assistants, adaptive surgical partners, and autonomous population health agents. However, they also emphasize that technical, ethical, and regulatory challenges must be addressed to ensure safe and effective clinical translation.

Radeva et al. [17] investigated awareness, trust, and expectations of AI for glaucoma care among Bulgarian ophthalmologists. The cross‐sectional survey included 156 ophthalmologists and residents and examined the influence of demographic factors such as age, gender, and professional experience. The study found varying levels of AI awareness, with less experienced clinicians being more informed. Trust in AI diagnosis and treatment was low, although younger respondents were more optimistic about the impact of AI. Qualitative responses highlighted diagnostic utility but also concerns about training deficiencies. The authors concluded that Bulgarian ophthalmologists show cautious optimism toward AI in glaucoma care and that targeted training is needed to build trust and support equitable digital health adoption.

Jin et al. [18] systematically reviewed multimodal AI in ophthalmology, focusing on applications, technical characteristics, and clinical value. Their review followed PRISMA guidelines and examined literature published between 2018 and 2025 across PubMed, Web of Science, Scopus, and Google Scholar. The included studies covered glaucoma, AMD, corneal diseases, ophthalmic emergency triage, cognitive impairment screening, diabetes complication screening, and chatbot‐based ophthalmic consultation. The authors reported that multimodal systems generally outperformed unimodal systems, with AUC improvements of 4%–5% and accuracy improvements of 2%–7%. They concluded that multimodal AI has broad prospects in ophthalmology, but future research should emphasize clinical validation, novel fusion strategies, interpretability, and lightweight models to support real‐world translation.

Ran et al. [19] examined the acceptance of ophthalmic AI for eye diseases through a literature review and qualitative analysis. Their study evaluated 16 eligible studies using a psychological model based on performance expectancy, effort expectancy, social influence, facilitating conditions, and relevant regulating factors such as gender, age, experience, and voluntariness of use. The authors found that most prior studies focused mainly on performance expectancy and effort expectancy, while social influence, facilitating conditions, and regulating factors were less deeply explored. They reported generally high acceptance of ophthalmic AI among patients, ophthalmologists, other professionals, and the general population, but also highlighted concerns regarding economic burden, privacy, model safety, trustworthiness, public awareness, accountability, cost‐effectiveness, clinician training, and the need for evidence‐based validation.

Chen and Bai [20] reviewed AI technology in ophthalmology public health, emphasizing screening, monitoring, risk prediction, resource allocation, health education, telemedicine, and patient management. Their review argues that AI can improve the quality and efficiency of ophthalmic public health, particularly for cataracts, DR, glaucoma, and myopia. At the same time, the authors identify major barriers to implementation, including interoperability with electronic health records (EHRs), data security and privacy, data quality, bias, algorithmic transparency, ethical concerns, and regulatory frameworks. Their work is especially relevant for understanding AI as a public‐health infrastructure tool rather than only as a disease‐classification technology.

The present review differs from the above works in its scope, organizing logic, and translational emphasis. Several recent reviews focus on specific subdomains, such as ophthalmic conversational AI and VQA [6], multimodal AI [7, 18], glaucoma diagnosis or management [8, 15], agentic AI [9, 16], corneal nerve imaging [10], clinician and patient perceptions [11, 17], end‐user acceptance of ophthalmic AI [19], optometric education [12], foundation models [13], ophthalmic public health [20], or precision medicine across multiple biomedical domains [14]. In contrast, the present article provides a broader clinically grounded synthesis of AI across major ophthalmic diseases, imaging modalities, methodological paradigms, evaluation practices, translational barriers, and emerging directions.

More importantly, the novelty of this review is not limited to broader topic coverage. It organizes ophthalmic AI through an integrated clinical and translational framework. Specifically, it combines a clinical‐use‐case‐oriented structure, evidence maturity and translational readiness mapping, threshold‐aware interpretation of evaluation metrics, scenario‐based deployment requirements, and cautious appraisal of foundation models, generative AI, multimodal AI, and edge AI. This enables the review to move beyond model performance summaries and instead examine what level of evidence is required before ophthalmic AI systems can be responsibly implemented in real‐world care. Thus, while existing works provide valuable insights into selected technologies, diseases, or stakeholder perspectives, the present review aims to provide a more comprehensive and clinically prioritized account of ophthalmic AI readiness, limitations, and deployment requirements. Table 1 compares latest literature against this study.

2.3 Distinctive Contribution of This Review

The distinctive contribution of this review lies not merely in covering AI applications in ophthalmology, but in organizing the field through an integrated clinical and translational framework. Many existing reviews summarize ophthalmic AI according to disease categories, imaging modalities, or algorithmic approaches. In contrast, this review emphasizes how AI systems may function within real ophthalmic care pathways and what level of evidence is required before such systems can be considered clinically mature.

• First, the review adopts a clinical‐use‐case‐oriented organization that connects AI methods with practical ophthalmic needs, including population screening, referral triage, disease grading, progression monitoring, prognosis, treatment guidance, workflow support, and automated reporting. This structure helps move the discussion beyond model performance alone and places AI outputs within the context of actual clinical decision‐making.

• Second, the review explicitly maps evidence maturity and translational readiness across major ophthalmic applications. By distinguishing between technically promising tasks and clinically deployable systems, it highlights that diabetic retinopathy [21, 22] screening, glaucoma [2325] progression assessment, AMD monitoring, cataract grading [26], infectious keratitis classification, keratoconus detection, foundation‐model applications, and edge‐based screening differ substantially in their current readiness for real‐world implementation.

• Third, the review provides a threshold‐aware interpretation of evaluation metrics. Rather than treating accuracy, sensitivity, specificity, receiver operating characteristic‐area under the curve (ROC‐AUC), precision‐recall‐area under the curve (PR‐AUC), calibration, and segmentation metrics as purely mathematical quantities, it interprets them in relation to missed disease, false referral burden, disease prevalence, clinical thresholds, uncertainty, and workflow consequences. This emphasis is important because a model with strong aggregate performance may still be unsafe or inefficient if its operating point is poorly matched to the intended clinical setting.

• Fourth, the review introduces scenario‐based deployment requirements for specific ophthalmic contexts, including community DR screening, glaucoma follow‐up, AMD monitoring, and anterior segment triage. This provides a practical bridge between algorithmic validation and clinical implementation by identifying the operational requirements needed for successful deployment, such as image quality control, referral linkage, longitudinal comparison, clinician override, uncertainty reporting, and confirmatory diagnostic pathways.

• Finally, this review offers a cautious appraisal of emerging paradigms such as foundation models, generative AI, multimodal AI, and edge AI. While these approaches may expand the scope of ophthalmic AI, their current clinical evidence remains uneven. The review therefore emphasizes hallucination risk, clinical grounding, accountability, robustness, external validation, and the gap between benchmark performance and deployment readiness.

3 Foundations of Artificial Intelligence in Ophthalmology

The development of AI in ophthalmology reflects a broader evolution in computational medicine, transitioning from handcrafted analytical pipelines to data‐driven, representation‐learning paradigms. Each methodological phase has contributed distinct strengths, shaping current systems that are increasingly capable of addressing complex clinical tasks across diverse imaging modalities and patient populations.

Figure 3 provides a conceptual taxonomy of AI paradigms that underpin modern ophthalmic AI systems. At the highest level, AI encompasses a broad set of computational techniques designed to emulate intelligent decision‐making, within which ML represents a data‐driven subset that enables models to learn patterns from clinical and imaging data. DL, as a further specialization of ML, plays a dominant role in ophthalmology due to its ability to extract hierarchical features from high‐dimensional inputs such as fundus photographs, OCT scans, and slit‐lamp images. The overlap between DL and NLP gives rise to advanced capabilities such as conversational AI and LLMs, which extend ophthalmic applications beyond image analysis to include clinical report generation, decision support, and patient interaction. In this context, DL‐based vision models primarily support tasks such as disease detection, segmentation, and progression analysis, while NLP‐driven components facilitate integration with EHRs and clinical workflows. Thus, the figure highlights the layered and intersecting nature of AI technologies, illustrating how multiple paradigms collectively contribute to a comprehensive and multimodal ophthalmic intelligence framework.

3.1 Terminological Boundaries

Given the rapid expansion of AI research in ophthalmology, precise use of terminology is essential. Several terms, including AI, ML, DL, transformers, foundation models, generative AI, and multimodal AI, are sometimes used interchangeably in the literature, although they refer to different levels of computational abstraction. A clear distinction among these concepts helps avoid conceptual ambiguity and supports a more accurate interpretation of ophthalmic AI studies.

• AI is the broadest term and refers to computational systems designed to perform tasks that normally require human intelligence, such as perception, classification, reasoning, prediction, and decision support. In ophthalmology, AI may therefore include systems for automated disease screening, image interpretation, risk stratification, clinical triage, and workflow assistance.

• ML is a subset of AI in which algorithms learn patterns from data rather than relying entirely on explicitly programmed rules. Traditional ML methods usually depend on manually extracted or engineered features, such as vessel caliber, lesion texture, optic disc morphology, corneal [27, 28] curvature indices, or VF parameters. These features are then used by classifiers or regression models to support diagnosis, grading, or prediction.

• DL is a specialized branch of ML based on multi‐layer neural networks that learn hierarchical representations directly from raw or minimally processed data. In ophthalmic imaging, DL has become especially influential because it can automatically identify clinically meaningful patterns in fundus [29] photographs, OCT scans, slit‐lamp images, and corneal tomography without requiring extensive handcrafted feature design.

• CNNs are a major class of DL models designed primarily for image analysis. They use convolutional filters to capture local spatial patterns such as edges, textures, lesions, retinal vessels, optic disc boundaries, and retinal layer structures. CNN‐based architectures have been widely used in ophthalmology for DR grading, glaucoma detection, AMD classification, cataract [30] assessment, keratitis recognition, and keratoconus screening.

• Transformers are attention‐based architectures that model relationships among different parts of the input using self‐attention mechanisms. Unlike CNNs, which emphasize local spatial feature extraction, transformers can capture long‐range dependencies across image regions or multimodal inputs. In ophthalmology, transformer‐based models are increasingly explored for fundus image analysis, OCT interpretation, multimodal fusion, and vision‐language applications.

• Foundation models refer to large‐scale pretrained models developed using broad and diverse datasets, which can subsequently be adapted to multiple downstream tasks. Their relevance in ophthalmology lies in their potential to support few‐shot learning, cross‐dataset transfer, automated reporting, and general‐purpose representation learning across imaging modalities. However, their clinical use requires careful validation because strong pretraining performance does not necessarily guarantee safety or reliability in real‐world ophthalmic settings. Generative AI denotes models capable of producing new content, such as text, images, synthetic data, or structured clinical reports. In ophthalmology, generative AI may support report generation, patient communication, dataset augmentation, or image synthesis. Nevertheless, generative outputs must be interpreted cautiously because hallucinated findings, unsupported recommendations, or clinically ungrounded explanations may create safety risks.

• Multimodal AI refers to systems that integrate information from multiple sources, such as fundus images, OCT volumes, OCTA, VF tests, EHRs, demographic variables, and genomic data. This approach is particularly important in ophthalmology because clinical decisions often depend on the combined interpretation of structural, functional, and patient‐specific information rather than a single imaging modality.

3.2 Machine Learning Approaches

Early applications of AI in ophthalmology were grounded in classical ML frameworks, where domain expertise played a central role in feature engineering. Researchers designed pipelines that explicitly extracted clinically meaningful descriptors, such as vessel caliber, tortuosity, optic disc boundaries, and textural patterns associated with retinal lesions. These features were then supplied to statistical classifiers, including support vector machines, random forests, logistic regression, and k‐nearest neighbors. This paradigm offered interpretability and control, as each step—from feature extraction to classification—was transparent and often aligned with clinical reasoning. For instance, quantifying the cup‐to‐disc ratio provided a direct proxy for glaucoma assessment, while microaneurysm detection informed DR grading. However, the reliance on handcrafted features imposed inherent limitations. Feature design required substantial domain knowledge, was often task‐specific, and struggled to generalize across imaging devices or patient populations. Subtle patterns that were not explicitly encoded in the feature set frequently remained undetected, constraining diagnostic performance. Despite these limitations, classical ML established foundational principles for ophthalmic AI, including the importance of data preprocessing, feature robustness, and validation strategies [31].

3.3 Deep Learning Approaches

The advent of DL marked a decisive shift from manual feature engineering to automated representation learning. CNNs, in particular, transformed ophthalmic image analysis by enabling models to learn hierarchical features directly from raw pixel data. Lower layers capture basic structures such as edges and textures, while deeper layers encode increasingly abstract and task‐relevant representations. This capability proved especially effective for retinal imaging, where subtle variations in color, texture, and morphology can carry significant diagnostic meaning. Architectures such as ResNet and DenseNet introduced deeper networks with improved gradient flow, while EfficientNet emphasized parameter efficiency without sacrificing performance. U‐Net and its variants became the standard for segmentation tasks, enabling precise delineation of anatomical structures such as retinal layers, optic discs, and pathological regions. DL models also facilitated end‐to‐end optimization, reducing the need for intermediate processing steps and allowing direct mapping from input images to clinical outputs. This streamlined workflow contributed to substantial improvements in diagnostic accuracy across a range of conditions, including DR, AMD, and glaucoma [32].

3.4 Vision Transformers

The introduction of vision transformers has expanded the design space of DL models beyond convolutional architectures. Unlike CNNs, which rely on localized receptive fields and hierarchical feature aggregation, vision transformers model global dependencies through self‐attention mechanisms. By representing an image as a sequence of patches, these models enable direct interaction between distant regions, allowing the network to capture long‐range contextual relationships. In ophthalmology, where spatial distribution of pathological features often carries diagnostic significance, this global modeling capability is particularly advantageous. For example, the spatial arrangement of retinal lesions, diffuse structural changes in OCT volumes, or subtle asymmetries across anatomical regions may influence disease grading and progression assessment. Transformer‐based models can effectively encode such relationships without relying on progressively stacked convolutional filters [33]. In medical imaging domains where labeled data are limited, their performance may degrade without appropriate pretraining or regularization strategies. Consequently, their adoption in ophthalmology is often coupled with transfer learning or hybrid designs that mitigate data and resource constraints.

3.5 Multimodal Models

Clinical decision‐making in ophthalmology inherently involves multiple sources of information, including imaging, functional assessments, and patient‐specific factors. Multimodal models aim to replicate this integrative reasoning by learning joint representations from heterogeneous data sources, such as fundus images, OCT scans, VF measurements, EHRs, and demographic information [34]. By combining complementary modalities, these models can capture relationships that are not apparent within any single data source. For instance, integrating structural information from OCT with functional deficits observed in VFs can improve the detection and monitoring of glaucoma progression. Similarly, incorporating clinical history and risk factors enhances predictive modeling for diseases such as AMD.

3.6 Foundation Models

Foundation models represent a paradigm shift toward large‐scale, general‐purpose AI systems that can be adapted to a wide range of ophthalmic tasks. Trained on diverse and extensive datasets, these models learn rich and transferable representations that extend beyond narrow task‐specific objectives. As a result, they can be fine‐tuned for downstream applications with relatively limited labeled data. In ophthalmology, foundation models offer several advantages. They enable few‐shot and zero‐shot learning, which is particularly valuable for rare diseases or underrepresented populations. Their ability to generalize across imaging devices and acquisition conditions enhances robustness in real‐world deployment. Furthermore, when combined with language modeling capabilities, they support advanced applications such as automated report generation, clinical summarization, and decision support [35].

3.7 Explainable Artificial Intelligence

As AI systems become increasingly integrated into clinical workflows, the need for transparency and interpretability has gained prominence. In ophthalmology, where diagnostic decisions can have profound consequences for patient outcomes, clinicians must be able to understand and trust model outputs. XAI seeks to address this requirement by providing insights into the decision‐making process of complex models. Techniques such as saliency maps and Gradient‐weighted Class Activation Mapping (Grad‐CAM) highlight regions of an image that contribute most strongly to a prediction, offering a visual explanation that can be compared with clinical reasoning. Methods based on feature attribution, such as SHAP, quantify the contribution of individual inputs, while attention visualization provides an alternative perspective in transformer‐based models [36].

3.8 Public Ophthalmic Datasets

The advancement of AI in ophthalmology has been driven not only by methodological innovations but also by the availability of diverse, well‐annotated public datasets. These datasets span multiple imaging modalities, disease domains, and annotation strategies, enabling systematic benchmarking of classification, segmentation, grading, multimodal learning, and clinical decision support tasks. Table 2 summarizes representative datasets that have become foundational in ophthalmic AI research. Among color fundus photography datasets, several resources have been central to DR research. EyePACS [37] remains one of the most widely used benchmarks for DR screening and grading, particularly in large‐scale DL studies. APTOS 2019 [38] complements this by providing data acquired under more variable screening conditions, supporting robustness evaluation. MESSIDOR [39] has long served as a reference dataset for computer‐assisted DR diagnosis, while MESSIDOR‐2 [40] is frequently used for external validation, despite the absence of publicly available ground‐truth labels. IDRiD [41] extends these datasets by providing detailed lesion‐level annotations alongside DR and diabetic macular edema (DME) grading, enabling multi‐task learning and fine‐grained segmentation studies.

Beyond DR, several fundus‐based datasets support broader disease analysis. ODIR‐2019 (ODIR‐5K) [42] enables multi‐label classification across multiple ocular conditions, incorporating demographic metadata alongside bilateral fundus images. Disease‐specific datasets such as ADAM [43] and PALM [44] facilitate research in AMD and pathological myopia, respectively, with annotations supporting both classification and structural analysis. In glaucoma research, REFUGE [45] provides clinically annotated labels with optic disc and cup segmentation, while ORIGA‐light [46] and RIGA [47] contribute additional resources for cup‐to‐disc ratio estimation and inter‐observer variability analysis.

The dataset landscape further expands with multimodal and cross‐sectional imaging. GAMMA [48] integrates fundus photography with three‐dimensional OCT data to support multimodal glaucoma grading, reflecting the increasing importance of cross‐modal learning. RETOUCH [49] serves as a dedicated OCT benchmark for retinal fluid detection and segmentation, with annotations for clinically relevant fluid types. Similarly, the Duke DME OCT dataset [50] provides expert‐labeled retinal layers and fluid regions, supporting detailed structural analysis in DME.

For vascular analysis, several classical fundus datasets remain essential. DRIVE [51], STARE [52], and CHASE_DB1 [53] provide manually annotated vessel masks, enabling supervised vessel segmentation and evaluation of model robustness. RITE [54] extends this line of work by supporting artery–vein classification and vessel tree analysis. In the context of OCTA, ROSE [55] introduces a dedicated benchmark for vascular segmentation, with annotations at both centerline and pixel levels.

4 Ophthalmic Imaging Modalities

The success of AI in ophthalmology is fundamentally anchored in the richness and diversity of imaging modalities. Each modality captures distinct anatomical or functional aspects of the eye, shaping both the design of AI models and the clinical questions they are intended to address. Beyond the imaging technologies themselves, the surrounding data ecosystem—including datasets, annotations, and curation practices—plays a decisive role in determining the reliability and generalizability of AI systems.

4.1 Fundus Photography

Fundus photography has long served as the cornerstone of ophthalmic imaging and remains the most extensively utilized modality in AI‐driven applications. Its widespread adoption is largely due to its accessibility, relatively low cost, and ability to capture a comprehensive view of the retinal surface in a single image. From an algorithmic perspective, fundus images provide a rich visual landscape for detecting vascular abnormalities, lesions, and structural changes. This has made the modality particularly suitable for large‐scale screening tasks, such as DR detection, glaucoma risk assessment through optic disc analysis, and identification of hypertensive retinal changes. The availability of large, annotated datasets has further accelerated progress, enabling the development of highly performant models [56]. At the same time, fundus photography presents notable challenges. Variability in image quality, differences in camera specifications, and the presence of artifacts such as glare or poor focus can significantly influence model performance. As such, robust preprocessing and quality assessment mechanisms are often essential components of fundus‐based AI systems.

4.2 Optical Coherence Tomography

OCT represents a major advancement in retinal imaging, providing high‐resolution cross‐sectional views of retinal layers. Unlike fundus photography, which offers a two‐dimensional projection, OCT enables detailed visualization of subsurface structures, making it indispensable for the diagnosis and monitoring of macular diseases. AI models applied to OCT data are typically designed to analyze volumetric information, capturing subtle variations in retinal thickness, fluid accumulation, and layer integrity. This is particularly relevant in conditions such as AMD and DME, where structural changes evolve over time. The ability to process three‐dimensional data introduces additional computational complexity but also opens the door to more precise and clinically meaningful analysis [57]. One of the distinguishing features of OCT‐based AI is its role in longitudinal monitoring. Repeated scans allow models to detect progression patterns that may not be evident in isolated images, thereby supporting more informed treatment decisions.

4.3 OCT Angiography

OCTA extends the capabilities of conventional OCT by enabling non‐invasive visualization of retinal and choroidal microvasculature. By capturing motion contrast from blood flow, OCTA provides insights into vascular integrity without the need for dye injection, reducing patient risk and procedural complexity. In the context of AI, OCTA data introduce a different analytical focus centered on vascular patterns. Models are often tasked with quantifying vessel density, identifying areas of non‐perfusion, and detecting abnormal neovascularization. These features are particularly relevant in diseases such as DR and retinal vein occlusion, where microvascular changes play a central role [58]. However, OCTA data are inherently sensitive to motion artifacts and noise, which can complicate analysis. Moreover, the lack of standardized acquisition protocols across devices poses challenges for model generalization. Addressing these issues requires careful preprocessing and cross‐device validation.

4.4 Slit‐Lamp and Anterior Segment Imaging

While much of ophthalmic AI research has focused on posterior segment imaging, anterior segment modalities such as slit‐lamp photography and corneal topography are equally important in clinical practice. These modalities capture the optical and structural properties of the cornea, lens, and ocular surface. AI applications in this domain often involve texture and pattern analysis, which differ fundamentally from the vascular and structural features observed in retinal imaging. Tasks such as cataract grading rely on assessing lens opacity, whereas keratoconus detection depends on identifying subtle distortions in corneal curvature. Similarly, the classification of corneal infections requires distinguishing between visually similar but clinically distinct patterns [59]. The diversity of anterior segment conditions and imaging techniques introduces variability that can challenge model robustness. Consequently, successful systems must be adaptable to different acquisition settings while maintaining sensitivity to clinically relevant features.

4.5 Visual Fields and Functional Data

Not all ophthalmic information is derived from imaging. Functional assessments, such as VF testing, provide critical insights into how structural changes translate into vision loss. These data are inherently temporal and often exhibit high variability, reflecting both disease progression and patient‐specific factors. AI approaches applied to functional data focus on pattern recognition over time, identifying trends that may indicate progression or stability. In glaucoma, for instance, correlating VF deterioration with structural changes observed in OCT can enhance diagnostic confidence and improve prognostic accuracy. Unlike image‐based modalities, functional data require models capable of handling sequential inputs and irregular sampling intervals [60]. This introduces a different set of methodological challenges, emphasizing temporal modeling and robustness to missing or noisy data. Table 3 compares major ophthalmic imaging modalities in AI.

5 Clinical Use‐Case‐Oriented View of Ophthalmic AI

Although ophthalmic AI is often discussed according to disease categories or model architectures, its clinical value is best understood through specific use cases. In practice, AI systems are expected to support defined clinical functions, such as screening, triage, grading, monitoring, prognosis, treatment planning, and workflow assistance. A use‐case‐oriented view therefore helps connect computational performance with real ophthalmic decision‐making. It also clarifies that the same algorithmic output may have different clinical implications depending on whether the system is used in a community screening program, a specialist clinic, a teleophthalmology service, or a longitudinal follow‐up pathway. Figure 4 shows how AI systems support population screening, referral triage, disease grading, progression monitoring, treatment planning, and workflow support across ophthalmic practice.

5.1 Population Screening

Population screening is one of the most mature and clinically relevant applications of AI in ophthalmology. Screening programs aim to identify individuals with referable or vision‐threatening disease before irreversible visual loss occurs. In this context, AI can assist by automatically analyzing fundus photographs, OCT scans, or anterior segment images and classifying cases according to referral need. DR screening represents the strongest example of this use case. AI systems can support large‐scale screening by detecting referable DR, DME, or ungradable images from fundus photographs. Similar screening‐oriented applications are emerging for glaucoma risk assessment, cataract detection, AMD identification, and keratoconus screening [61]. In low‐resource or community settings, AI may help extend specialist‐level assessment to primary care centers, mobile eye camps, and teleophthalmology networks. For screening applications, sensitivity is usually prioritized because missed disease can delay referral and treatment. However, specificity remains important because excessive false positives may overload ophthalmology clinics and reduce the efficiency of screening programs. Therefore, successful screening systems require not only high diagnostic performance but also image‐quality assessment, clear referral thresholds, local workflow integration, and mechanisms for follow‐up completion.

5.2 Referral Triage and Prioritization

Referral triage involves categorizing patients according to urgency, severity, or need for specialist review. This use case is distinct from general screening because the goal is not merely to detect disease, but to prioritize clinical attention. AI‐based triage systems can help separate normal or low‐risk cases from urgent, referable, or high‐risk cases, thereby improving the allocation of limited ophthalmic resources. In DR, triage may involve distinguishing non‐referable disease from referable or vision‐threatening disease. In AMD, AI can identify OCT features such as subretinal or intraretinal fluid that may require timely retinal specialist review. In glaucoma, triage may involve identifying patients with suspicious optic nerve features, progressive retinal nerve fiber layer loss, or abnormal VFs. In anterior segment disease, AI may assist in flagging suspected microbial keratitis, corneal ulceration, severe cataract, or other conditions requiring urgent evaluation. The clinical safety of triage systems depends on transparent threshold selection, uncertainty reporting, and escalation pathways. Borderline, poor‐quality, or out‐of‐distribution cases should not be forced into definitive categories without human review. Thus, AI‐based triage should function as a risk‐prioritization tool that supports clinical decision‐making while preserving clinician oversight.

5.3 Disease Grading and Severity Stratification

Disease grading and severity stratification are central to ophthalmic practice because management decisions often depend on disease stage rather than simple disease presence or absence. AI systems can support this process by assigning severity grades, quantifying lesions, measuring anatomical structures, or estimating clinically relevant biomarkers. In DR, AI can classify disease severity according to established grading scales and identify features such as microaneurysms, hemorrhages, exudates, and neovascularization. In cataract, AI can assist in grading lens opacity and estimating visual significance from slit‐lamp or fundus images. In keratoconus, AI can support severity stratification using corneal topography, tomography, pachymetry, and biomechanical parameters. In glaucoma, severity assessment may involve optic disc analysis, retinal nerve fiber layer thickness, VF loss, and cup‐to‐disc ratio estimation. Similarly, AMD grading may require detection of drusen, pigmentary changes, geographic atrophy, or neovascular features. The clinical usefulness of grading systems depends on consistency, interpretability, and alignment with accepted clinical classification schemes. A model that assigns a severity label without explaining the anatomical or pathological basis of that label may have limited practical value. Therefore, grading‐oriented AI should ideally provide supporting evidence, such as lesion localization, biomarker quantification, structural measurements, or confidence estimates.

5.4 Progression Monitoring and Prognosis

Many ophthalmic diseases require longitudinal monitoring rather than one‐time diagnosis. Progression monitoring is especially important in chronic conditions such as glaucoma, AMD, DME, keratoconus, and inherited retinal diseases. AI systems can assist by comparing current and previous examinations, identifying subtle changes, and estimating future risk. In glaucoma, progression assessment may require integration of OCT‐derived structural parameters with VF trends over time. AI models can help detect retinal nerve fiber layer thinning, optic nerve head changes, or functional deterioration that may not be obvious from isolated measurements. In AMD, longitudinal OCT analysis can support detection of fluid recurrence, geographic atrophy expansion, or conversion to neovascular disease. In keratoconus, serial topography or tomography can be used to identify progressive corneal steepening or thinning, thereby informing the timing of interventions such as corneal collagen cross‐linking. Prognostic AI models aim to estimate future disease trajectory, treatment response, or risk of visual decline. However, reliable prognostic modeling requires longitudinal datasets, standardized follow‐up intervals, and careful handling of missing or irregularly sampled data. Since disease progression may be slow, heterogeneous, and influenced by treatment, prognostic AI must be validated prospectively before being used to guide major clinical decisions.

5.5 Treatment Guidance and Follow‐Up Planning

Beyond diagnosis and monitoring, AI has potential to support treatment guidance and follow‐up planning. In this use case, the model output is intended to inform clinical management, such as referral urgency, treatment initiation, retreatment decisions, or follow‐up interval selection. This is a higher‐risk application than screening or grading because it may directly influence patient care pathways. In AMD and DME, OCT‐based AI systems may help quantify retinal fluid, estimate disease activity, and support anti‐vascular endothelial growth factor treatment planning. In cataract, AI may assist in surgical prioritization by combining lens opacity grading, visual function, and patient‐specific factors. In keratoconus, AI may help identify patients at risk of progression and support decisions regarding corneal collagen cross‐linking. In infectious keratitis, AI may assist early recognition of suspicious patterns, but definitive treatment guidance must remain linked to clinical examination and microbiological confirmation. Treatment‐support systems require particularly strong validation because incorrect recommendations may lead to overtreatment, undertreatment, or delayed intervention. Therefore, AI should be positioned as an assistive tool that provides quantitative evidence, risk estimates, or decision support, while final treatment decisions remain under clinician responsibility. Clear uncertainty communication, explainability, and auditability are essential in this context.

5.6 Workflow Support and Automated Reporting

AI can also contribute to ophthalmology by improving workflow efficiency rather than directly replacing diagnostic judgment. Workflow‐oriented applications include image quality assessment, automated measurement, structured reporting, prioritization of review queues, summarization of longitudinal findings, and integration with teleophthalmology or EHR systems. Automated reporting is an emerging application, especially with the development of multimodal and vision‐language models. Such systems may generate preliminary descriptions of fundus images, OCT findings, lesion burden, or disease activity. They may also assist clinicians by summarizing prior visits, highlighting interval changes, and suggesting items requiring attention. In high‐volume clinics, these tools may reduce documentation burden and improve consistency of reporting. However, automated reports must be carefully grounded in image evidence and clinical context. Generative outputs may contain unsupported statements, omitted findings, or hallucinated interpretations. Therefore, report‐generation systems should be designed with clinician review, traceable image evidence, structured templates, and safeguards against unsupported recommendations. In workflow support, the goal of AI should be to reduce repetitive burden, improve consistency, and enhance clinical efficiency without weakening accountability.

6 Key Disease‐Specific AI Applications

The impact of AI in ophthalmology becomes most evident when examined through specific disease contexts. Each condition presents unique diagnostic challenges, imaging characteristics, and clinical priorities, shaping the way AI systems are developed and evaluated. Rather than a uniform application of algorithms, progress in this domain reflects a series of tailored solutions aligned with disease‐specific needs. Figure 5 shows disease‐specific mapping of ophthalmic AI according to dominant modality, major task, and clinical role.

6.1 Diabetic Retinopathy

DR represents one of the most mature and clinically validated applications of AI in ophthalmology, owing to well‐defined grading standards, large annotated datasets, and the global need for scalable screening. AI systems have progressed from lesion‐level detection to end‐to‐end pipelines capable of identifying referable disease, grading severity, and supporting referral decisions.

Early large‐scale studies demonstrated the feasibility of automated DR detection using deep convolutional architectures. For instance, Ting et al. [62] and Gulshan et al. [63] reported high diagnostic performance (AUC up to 0.99) for referable DR using fundus images, establishing a benchmark for subsequent systems. Similar approaches have been extended to detect vision‐threatening DR and DME, with models such as EfficientNet‐based frameworks achieving robust performance across large datasets [64, 65].

Beyond classification, more advanced models incorporate grading and lesion‐level analysis. Approaches based on multiple‐instance learning [66] and segmentation‐assisted architectures [67] enable finer characterization of disease severity. Ensemble strategies and commercial systems further high‐light variability in real‐world performance [68], underscoring the importance of external validation.

Notably, AI‐driven screening has expanded to resource‐constrained settings through smartphone‐based fundus imaging, where systems such as the Medios platform demonstrate high sensitivity with acceptable specificity [69]. However, challenges remain in ensuring generalizability across populations, imaging conditions, and device variability.

6.2 Glaucoma

Glaucoma is characterized by its insidious onset and progressive neurodegeneration, making early detection and longitudinal monitoring essential. AI research in this domain spans multiple modalities, including fundus imaging for optic disc assessment, OCT for retinal nerve fiber layer evaluation, and VF testing for functional analysis.

A key strength of glaucoma‐related AI lies in its ability to integrate structural and functional information. Multimodal approaches combining OCT and VF data have demonstrated improved diagnostic performance, as evidenced by fusion‐based frameworks [70] and longitudinal progression models using convolutional LSTM networks [71]. Similarly, hybrid systems leveraging fundus images and clinical data enable prediction of disease onset and progression with high reliability [72]. These approaches are particularly valuable, as structural changes often precede measurable functional loss.

DL architectures have also achieved strong performance in modality‐specific tasks. OCT‐based systems, including 3D convolutional networks, show robust capability in detecting glaucomatous optic neuropathy [73], while fundus‐based models such as Inception and ResNet variants achieve high accuracy for referable glaucoma detection [64, 74]. Additionally, anterior‐segment OCT‐based models extend AI applications to angle‐closure assessment [75], highlighting the breadth of imaging‐driven approaches.

Beyond supervised learning, unsupervised and ML methods contribute to functional monitoring and disease characterization. For example, clustering‐based approaches applied to VF data enable detection of subtle functional deterioration [76], while traditional ML models incorporating intraocular pressure and sensor‐derived parameters provide complementary diagnostic insights [77].

6.3 Age‐Related Macular Degeneration

AMD presents unique challenges due to its heterogeneous structural manifestations and variable progression patterns. AI systems in this domain primarily utilize OCT and fundus imaging to identify key biomarkers such as drusen, pigmentary abnormalities, and intra‐/subretinal fluid.

DL models have demonstrated strong performance in both detection and staging tasks. Convolutional architectures applied to OCT and fundus images achieve high diagnostic accuracy, with studies reporting AUC values approaching 0.99 for differentiating disease stages and detecting referable AMD [78, 79]. Similarly, large‐scale fundus‐based models enable reliable classification of AMD severity and identification of neovascular AMD [80, 81], while ensemble approaches further enhance staging performance across diverse datasets [82].

Beyond static classification, there is increasing emphasis on modeling disease progression and risk prediction. Longitudinal analysis of OCT volumes, including 3D and temporal DL models, enables the identification of progression biomarkers and prediction of conversion to advanced stages such as wet AMD [8385]. These approaches are particularly relevant in clinical management, where treatment decisions—such as anti‐VEGF therapy—depend on dynamic structural changes over time.

Recent work also explores self‐supervised learning paradigms to improve generalization and reduce annotation dependency [86], addressing challenges associated with data variability. Nevertheless, the heterogeneity of disease presentation, differences in imaging protocols, and variability across populations continue to limit model robustness in real‐world settings.

6.4 Cataract

In contrast to retinal diseases, cataract primarily affects the optical transparency of the lens, leading to progressive visual impairment. AI applications in this domain focus on screening, severity grading, and, more recently, clinical decision support for treatment planning.

DL models applied to slit‐lamp and fundus images have demonstrated high diagnostic performance for cataract detection. For instance, ResNet‐based architectures achieve near‐perfect discrimination of cataract and referable cases [87], while attention‐based and hybrid models enable reliable grading of disease severity [88, 89]. Additional approaches combining convolutional networks with traditional ML methods further improve detection of visually significant cataracts [90]. These systems aim to standardize grading procedures that are otherwise subject to inter‐observer variability.

Beyond detection, AI has been extended to more specialized clinical tasks. Systems such as CC‐Cruiser support pediatric cataract diagnosis and treatment recommendation [91], while models incorporating demographic and clinical variables enable identification of congenital cataracts [92]. Furthermore, quantitative grading frameworks and regression‐based models provide continuous severity estimation, offering finer clinical interpretation [93, 94].

6.5 Infectious Keratitis

Infectious keratitis represents a clinically critical and time‐sensitive domain within ophthalmic AI, characterized by rapid progression and risk of irreversible vision loss. Unlike chronic retinal diseases, effective management requires early and accurate differentiation between causative pathogens, including bacterial, fungal, and viral etiologies.

AI systems in this domain leverage heterogeneous data sources such as slit‐lamp images, corneal photographs, and IVCM. DL models applied to slit‐lamp imagery demonstrate strong capability in detecting keratitis and distinguishing infectious from non‐infectious cases, with architectures such as DenseNet and Inception achieving high diagnostic performance [95, 96]. More fine‐grained classification, particularly differentiation between bacterial and fungal keratitis, has been addressed using ensemble learning and lightweight architectures, though performance remains variable due to overlapping visual features [97, 98].

Multiclass classification frameworks further extend this capability to distinguish multiple etiologies simultaneously, as demonstrated by large‐scale models trained on diverse datasets [99, 100]. Additionally, modality‐specific approaches such as IVCM‐based DL models achieve high accuracy for fungal keratitis diagnosis [101], while emerging non‐imaging approaches, including Raman spectroscopy and microbiome‐based analysis, highlight the potential of multimodal integration for improved diagnostic precision [102, 103].

6.6 Keratoconus

Keratoconus represents a clinically significant corneal disorder where early detection is both challenging and essential for preventing irreversible visual impairment. Characterized by progressive corneal thinning and protrusion, the disease often manifests through subtle structural alterations that are difficult to detect using routine examination, thereby creating a critical window for timely intervention such as corneal collagen cross‐linking.

AI‐based approaches in keratoconus primarily leverage corneal topography and tomography, which provide detailed representations of curvature, elevation, and thickness distribution. DL models applied to tomographic images have demonstrated high diagnostic performance, with architectures such as InceptionResNetV2 and EfficientNet‐based systems achieving AUC values up to 0.99 for detecting keratoconus and irregular corneal patterns [104, 105]. Similarly, models based on Pentacam‐derived parameters and biomechanical features enable accurate classification using both ML and hybrid approaches [106108].

A key advancement in this domain is the detection of subclinical keratoconus, which is critical for refractive surgery screening. AI systems trained on corneal topographic and tomographic data have shown strong capability in identifying early‐stage or forme fruste keratoconus, even when conventional clinical indicators are inconclusive [109, 110]. Neural network‐based models and alternative architectures further support detection of suspected cases and improve sensitivity to early geometric distortions [111]. Table 4 shows the disease specific mapping to AI. Table 5 presents the evidence maturity and translational readiness of AI applications in ophthalmology.

7 Evaluation Metrics

Robust evaluation is central to the safe and effective deployment of AI systems in ophthalmology. Unlike conventional computer vision tasks, medical AI but their clinical relevance depends on disease prevalence, workflow context, and decision threshold.

7.1 Conventional Metrics

Most ophthalmic AI studies report a standard set of classification metrics derived from the confusion matrix. Let TP, TN, FP, and FN denote true positives, true negatives, false positives, and false negatives, respectively. Based on these quantities, several widely used performance measures are defined.

1. Accuracy: Accuracy represents the overall proportion of correctly classified instances:

(1)Accuracy=TP+TNTP+TN+FP+FN.

While accuracy provides a global measure of correctness, it can be misleading in imbalanced datasets, which are common in ophthalmology (e.g., low prevalence of referable DR).

2. Sensitivity (Recall or true positive rate): Sensitivity measures the ability of a model to correctly identify diseased cases:

(2)Sensitivity=TPTP+FN.

In screening applications, such as DR detection, high sensitivity is crucial to minimize missed diagnoses that could lead to irreversible vision loss.

3. Specificity (True negative rate): Specificity quantifies the ability to correctly identify non‐diseased cases:

(3)Specificity=TNTN+FP.

High specificity reduces unnecessary referrals, thereby lowering healthcare burden and improving system efficiency.

4. Precision (Positive predictive value): Precision reflects the reliability of positive predictions:

(4)Precision=TPTP+FP.

In clinical workflows, precision is particularly important when false positives lead to costly or invasive follow‐up procedures.

5. Recall: Recall is mathematically identical to sensitivity:

(5)Recall=TPTP+FN.

The term “recall” is more commonly used in ML literature, whereas “sensitivity” is preferred in clinical contexts.

6. F1‐score: The F1‐score provides a harmonic mean of precision and recall:

(6)F1score=2Precision×RecallPrecision+Recall.

This metric is especially useful when balancing false positives and false negatives is important, as in automated screening systems.

7. Receiver operating characteristic and AUC: The ROC curve plots sensitivity against the false positive rate:

(7)FPR=FPFP+TN.

The ROC‐AUC summarizes the model's discriminative ability across all classification thresholds:

(8)ROCAUC=01TPR(x)dx.

A higher ROC‐AUC indicates better separability between diseased and non‐diseased classes. However, ROC‐AUC may overestimate performance in highly imbalanced datasets.

8. Precision‐recall curve and PR‐AUC: For imbalanced datasets, the PR curve is often more informative. It captures the trade‐off between precision and recall across thresholds. The area under this curve is given by:

(9)PRAUC=10Precision(r)dr,

where r denotes recall. PR‐AUC is particularly relevant in ophthalmology, where disease prevalence is typically low.

9. Clinical interpretation and limitations: Although these metrics are widely reported, their interpretation must be contextualized within clinical objectives. For example, a model with high accuracy but low sensitivity may be unsuitable for screening, while a model with high sensitivity but low precision may overwhelm healthcare systems with false referrals. Therefore, no single metric is sufficient; instead, a combination of metrics must be considered to assess both diagnostic performance and clinical utility.

Furthermore, many studies report only aggregate metrics without stratification across subgroups (e.g., age, ethnicity, device type), which can obscure performance disparities. As a result, there is a growing consensus that evaluation frameworks in ophthalmic AI should move beyond conventional metrics toward more holistic, reliability‐aware, and deployment‐oriented measures.

7.2 Segmentation Metrics

In ophthalmology, segmentation plays a critical role in tasks such as optic disc and cup delineation, retinal layer extraction in OCT, fluid detection, and lesion localization. Unlike classification, segmentation requires pixel‐level agreement between predicted masks and ground truth annotations, making evaluation more nuanced.

1. Dice similarity coefficient (DSC): The Dice similarity coefficient quantifies the overlap between predicted segmentation P and ground truth G:

(10)DSC=2|PG||P|+|G|.

The Dice score ranges from 0 to 1, where 1 indicates perfect overlap. It is widely used in medical imaging due to its robustness in handling class imbalance, particularly when the region of interest (e.g., lesions) occupies a small fraction of the image.

2. Intersection over union (IoU): Also known as the Jaccard index, IoU measures the ratio of overlap to union:

(11)IoU=|PG||PG|.

IoU is generally more stringent than Dice, as it penalizes mismatches more heavily. The relationship between Dice and IoU is given by:

(12)DSC=2IoU1+IoU.

3. Hausdorff distance (HD): While overlap‐based metrics evaluate region similarity, they do not capture boundary discrepancies. The Hausdorff distance measures the maximum boundary deviation:

(13)H(P,G)=maxsuppPinfgGd(p,g),supgGinfpPd(g,p),

where d(·, ·) denotes a distance metric (typically Euclidean). A lower Hausdorff distance indicates better boundary alignment. In practice, the 95th percentile Hausdorff distance (HD95) is often used to reduce sensitivity to outliers.

4. Boundary‐based metrics: To further evaluate contour accuracy, boundary‐based metrics such as average symmetric surface distance (ASSD) and contour matching scores are employed:

(14)ASSD(P,G)=1|P|+|G|pPmingGd(p,g)+gGminpPd(g,p).

These metrics are particularly important in applications like glaucoma, where small boundary deviations in optic cup segmentation can significantly affect clinical interpretation (e.g., cup‐to‐disc ratio).

5. Clinical considerations: From a clinical standpoint, segmentation errors near critical anatomical boundaries (e.g., fovea, optic nerve head) are often more consequential than uniform errors across the image. Therefore, combining region based and boundary‐based metrics provides a more comprehensive assessment of model performance.

7.3 Calibration and Reliability

In clinical settings, predictive accuracy alone is insufficient; models must also provide reliable confidence estimates. A well‐calibrated model produces probability outputs that reflect true likelihoods of correctness.

1. Brier score: The Brier score measures the mean squared error between predicted probabilities pi^ and true labels yi0,1:

(15)Brierscore=1Ni=1Npi^yi2.

Lower values indicate better calibration and overall probabilistic accuracy.

2. Expected calibration error (ECE): ECE quantifies the discrepancy between predicted confidence and empirical accuracy across M bins:

(16)ECE=m=1M|Bm|N|acc(Bm)confBm|,

where Bm denotes the set of samples in bin m. A lower ECE indicates better alignment between confidence and correctness.

3. Reliability diagrams: Reliability diagrams provide a visual representation of calibration by plotting predicted confidence against observed accuracy. Ideally, predictions should lie along the diagonal, indicating perfect calibration.

4. Uncertainty‐aware inference: Advanced methods incorporate epistemic and aleatoric uncertainty using techniques such as Monte Carlo dropout, deep ensembles, and Bayesian neural networks:

(17)Var(y|x)=Eθy2Eθ[y]2.

Uncertainty estimation enables risk‐aware decision‐making, allowing models to defer uncertain cases to clinicians.

5. Clinical relevance: In ophthalmology, miscalibrated models may produce overconfident predictions on poor‐quality images or out‐of‐distribution cases, potentially leading to harmful decisions. Therefore, calibration is essential for safe deployment.

7.4 Robustness and Generalization

For real‐world deployment, ophthalmic AI systems must generalize beyond the controlled conditions of training datasets. Robustness refers to consistent performance under varying conditions, while generalization refers to performance across unseen data distributions.

Domain shift: Domain shift arises when training and deployment data differ in distribution:

(18)Ptrain(x,y)Ptest(x,y).

This may result from differences in imaging devices, acquisition protocols, or patient populations.

2. Sources of variability: key sources of variability include:

◦ device heterogeneity (fundus camera types),

◦ demographic diversity (age, ethnicity),

◦ acquisition conditions (lighting, focus),

◦ image quality degradation (noise, blur),

◦ disease prevalence shifts.

3. Robustness evaluation: Robustness can be assessed through: (i) cross‐dataset validation, (ii) external multi‐center testing, (iii) noise and perturbation analysis, and (iv) adversarial robustness testing.

4. Generalization gap: The generalization gap can be expressed as:

(19)Δ=E(x,y)Ptest[l(f(x),y]E(x,y)Ptrain[l(f(x),y].

A large Δ indicates poor transferability to real‐world settings.

5. Clinical implications: Failure to generalize across devices or populations can lead to systematic bias and reduced diagnostic reliability, particularly in global health contexts where data heterogeneity is significant.

7.5 Clinical Utility Metrics

Beyond algorithmic performance, the ultimate goal of ophthalmic AI is to improve patient outcomes and healthcare efficiency. Therefore, evaluation must incorporate clinically meaningful utility metrics.

1. False referral burden: The proportion of non‐diseased cases incorrectly referred:

(20)Falsereferralrate=FPFP+TN.

High false referral rates can overload healthcare systems.

2. Missed disease burden: The proportion of diseased cases incorrectly classified as normal:

(21)Missrate=FNTP+FN,

This is critical in screening scenarios where missed diagnoses can lead to disease progression.

3. Operational efficiency: Metrics such as time savings and throughput improvement can be quantified as:

(22)Timegain=TmanualTAIassistedTmanual.

4. Cost‐effectiveness: Economic evaluation may consider cost per screened patient or cost per detected case:

(23)Costefficiency=TotalcostNumberofcorrectdiagnoses.

A consolidated comparison of evaluation metrics, their mathematical foundations, clinical interpretations, and limitations is presented in Table 6. The table highlights that no single metric is sufficient, emphasizing the need for a multi‐dimensional evaluation framework aligned with clinical objectives.

Although quantitative metrics are indispensable for evaluating ophthalmic AI systems, their clinical meaning depends strongly on the intended use case, disease prevalence, and decision threshold. A model that appears highly accurate in aggregate may still be unsuitable for clinical deployment if its errors occur in clinically harmful directions. Therefore, evaluation should not be limited to reporting numerical performance alone; it should also explain how each metric relates to screening safety, referral burden, diagnostic confidence, and downstream patient management.

In screening settings, sensitivity is often prioritized because missed disease may result in delayed diagnosis, irreversible visual loss, or loss to follow‐up. This is particularly important in conditions such as DR, glaucoma, and AMD, where early detection can substantially alter prognosis. A highly sensitive model reduces false‐negative decisions and therefore supports safer population‐level screening. However, sensitivity should not be interpreted in isolation, since extremely low thresholds may increase the number of false‐positive referrals.

Specificity becomes especially important when AI systems are deployed at scale. Low specificity may lead to excessive false referrals, unnecessary specialist consultations, patient anxiety, and inefficient use of limited ophthalmic resources. In community screening programs or teleophthalmology workflows, even a modest reduction in specificity can create a large operational burden when thousands of individuals are screened. Thus, an acceptable balance between sensitivity and specificity must be determined according to the clinical objective and available healthcare capacity.

Threshold‐independent measures, such as ROC‐AUC are useful for summarizing discriminative performance across all possible thresholds. However, a high ROC‐AUC does not necessarily imply strong performance at the specific operating threshold used in clinical practice. For example, two models may have similar ROC‐AUC values but differ substantially in the number of missed referable cases or false referrals at a clinically selected decision point. Hence, ROC‐AUC should be complemented by threshold‐specific sensitivity, specificity, positive predictive value, negative predictive value, and referral‐rate estimates.

In low‐prevalence screening settings, PR‐AUC may provide more clinically informative evidence than ROC‐AUC. Many ophthalmic screening tasks involve relatively few diseased or referable cases compared with normal cases. Under such class imbalance, ROC‐AUC may appear favorable even when the positive predictive value is modest. PR‐AUC gives greater emphasis to the model's ability to correctly identify positive cases and is therefore particularly relevant when the clinical concern is reliable detection of uncommon but vision threatening disease.

Calibration is another essential consideration when AI confidence scores are used for triage or decision support. A model may correctly rank cases by risk but still produce poorly calibrated probabilities. In clinical workflows, this distinction is important because predicted probabilities may influence urgency of referral, need for clinician review, or patient counseling. Overconfident predictions on poor‐quality images, rare disease presentations, or out‐of‐distribution cases may create unsafe reassurance or unnecessary escalation. Therefore, calibration measures such as reliability diagrams, Brier score, and expected calibration error should be interpreted alongside conventional diagnostic metrics.

For segmentation tasks, overlap‐based measures such as Dice similarity coefficient and intersection over union are useful but may not fully capture clinical adequacy. A high Dice score can coexist with boundary errors that are clinically meaningful, particularly when segmenting the optic cup, foveal region, retinal layers, fluid compartments, or corneal structures. Small contour deviations may alter derived biomarkers such as cup‐to‐disc ratio, retinal thickness, or lesion volume. Consequently, segmentation evaluation should combine overlap metrics with boundary‐sensitive measures and, where possible, clinically derived endpoint errors.

Ultimately, threshold choice should be disease‐specific and workflow‐specific rather than purely algorithmic. A DR screening system may require a threshold that minimizes missed referable disease, whereas a specialist referral triage system may require a threshold that balances urgency with clinic capacity. Similarly, glaucoma progression monitoring, AMD fluid detection, cataract grading, and infectious keratitis triage each involve different risks associated with false negatives and false positives. Therefore, ophthalmic AI studies should report not only global performance metrics but also the rationale for threshold selection, expected referral consequences, and the clinical context in which the chosen operating point is intended to be used.

8 Translational Challenges

Despite remarkable progress in algorithm development, the translation of ophthalmic AI from controlled research settings to routine clinical practice remains a complex and often underestimated endeavor. High benchmark performance does not automatically translate into clinical reliability, and several non‐technical factors—ranging from workflow compatibility to regulatory approval—play a decisive role in determining real‐world success. In practice, deployment is not a single step but a multi‐layered process that requires alignment between technology, clinicians, healthcare infrastructure, and policy frameworks.

8.1 Scenario‐Based Requirements for Clinical Deployment

The requirements for deploying AI in ophthalmology vary substantially according to the clinical scenario in which the system is used. A model that performs well in a retrospective dataset may not necessarily be suitable for real‐world practice unless it is aligned with the operational needs, safety requirements, and decision pathways of the intended setting. Therefore, translational readiness should be assessed through concrete deployment scenarios rather than through aggregate performance metrics alone. Figure 6 presents the scenario‐based deployment requirements for ophthalmic AI. Different clinical settings require distinct operational safeguards, including image quality control, longitudinal comparison, uncertainty reporting, clinician override, and confirmatory diagnostic pathways.

In community DR screening, AI systems must support high‐throughput assessment while minimizing missed referable disease. Since image acquisition is often performed outside specialist clinics, image quality control is essential. The system should be able to detect ungradable or poor‐quality fundus images and request repeat acquisition where necessary. Referral thresholds must be carefully selected to balance sensitivity against false referral burden. In addition, locally understandable reporting, including local language summaries where appropriate, can improve communication with patients and primary care workers. Uncertain or borderline cases should be routed for human review rather than automatically classified as normal or abnormal. Audit trails are also necessary to document image quality, model output, confidence score, referral recommendation, and final clinical decision. Finally, screening systems must be linked to follow‐up pathways, since detection without referral completion does not improve visual outcomes.

For glaucoma follow‐up, deployment requirements are different because the central clinical challenge is not only diagnosis but also longitudinal assessment of progression. AI systems should integrate structural data from OCT with functional information from VF testing. The model should compare current findings with prior visits, identify clinically meaningful change, and distinguish true progression from test variability or imaging artifacts. Progression alerts may assist clinicians by highlighting patients who require closer monitoring or treatment escalation. However, because glaucoma progression assessment often involves nuanced interpretation, clinician override must remain available. The system should therefore function as a decision‐support tool that enhances longitudinal surveillance rather than as an autonomous replacement for specialist judgment.

In AMD monitoring, AI deployment should emphasize quantitative assessment and treatment‐relevant interpretation. OCT‐based systems should be able to identify and quantify intraretinal fluid, subretinal fluid, pigment epithelial detachment, and other structural biomarkers that influence treatment decisions. Beyond detecting fluid, clinically useful systems should support treatment‐response prediction and assist in planning follow‐up intervals for anti‐vascular endothelial growth factor therapy. Because management decisions may depend on small anatomical changes over time, uncertainty reporting is essential. The system should clearly indicate when image quality, atypical anatomy, or borderline fluid detection reduces confidence. Such transparency can help clinicians decide whether to accept the AI output, repeat imaging, or perform additional review.

Anterior segment triage presents another distinct deployment context, especially for cataract, keratitis, corneal opacity, and ocular surface disorders. Successful implementation requires robust slit‐lamp or anterior segment image capture under variable lighting, magnification, and acquisition conditions. For infectious keratitis, AI systems should not merely identify corneal abnormalities but should help differentiate infectious from non‐infectious causes and, where possible, distinguish likely bacterial, fungal, viral, or acanthamoeba patterns. However, visual overlap among these conditions limits the reliability of image‐only diagnosis. Therefore, urgent referral flags should be incorporated for cases with suspected microbial keratitis, corneal ulceration, perforation risk, or rapidly progressive disease. AI output should also be linked with microbiology confirmation pathways, since culture, smear, polymerase chain reaction, or confocal microscopy may be required for definitive diagnosis and treatment selection.

8.2 Workflow Integration

The effectiveness of an AI system in ophthalmology is closely tied to how naturally it fits within existing clinical workflows. Systems that operate in isolation or require additional manual steps are unlikely to gain sustained acceptance, regardless of their technical accuracy. Instead, AI tools must function as an extension of routine practice—integrated seamlessly with EHRs, imaging systems such as picture archiving and communication system (PACS) and teleophthalmology platforms.

From an operational perspective, the introduction of AI should simplify rather than complicate the diagnostic process. Ideally, it should reduce the time required for image interpretation while preserving clinician oversight and confidence. A useful way to conceptualize this is through the total diagnostic pathway:

(24)Ttotal=Tacquisition+Tanalysis+Tclinicalreview,

where meaningful deployment seeks to optimize, rather than merely automate, each component.

Equally important is the human factor. Systems that generate excessive alerts, require frequent manual corrections, or provide opaque outputs can increase cognitive burden and lead to disengagement. Therefore, successful integration depends as much on usability, interface design, and clinician trust as it does on algorithmic performance.

8.3 Regulatory Considerations

In most jurisdictions, ophthalmic AI systems are classified as software as a medical device, subjecting them to rigorous regulatory scrutiny. Approval processes increasingly demand not only technical validation but also evidence of clinical benefit, reproducibility, and safety across diverse populations. A particular challenge arises with adaptive or continuously learning systems. Unlike static models, these systems may evolve over time as new data become available. While this adaptability offers potential improvements in performance, it also introduces uncertainty regarding consistency and accountability. Regulatory frameworks are therefore evolving to address questions such as: how should updates be validated, how can decision pathways remain traceable, and what mechanisms ensure that performance does not degrade in specific subgroups? Post‐deployment monitoring is equally critical. Real‐world performance may diverge from initial validation results, making ongoing auditing, feedback loops, and periodic re‐evaluation essential components of responsible deployment.

8.4 Data Privacy and Security

The use of sensitive patient data places stringent requirements on privacy and security. Ophthalmic images, often linked with clinical metadata, must be handled in accordance with established data protection standards. This extends beyond simple anonymization to include secure storage, controlled access, and protection against unauthorized inference or reconstruction. Emerging paradigms such as federated learning offer a promising direction by enabling collaborative model training without centralizing raw data. In such settings, learning occurs across distributed institutions, with only model updates shared. This approach reduces the risk of data exposure while allowing models to benefit from diverse datasets. However, privacy‐preserving methods are not without trade‐offs. They may introduce communication overhead, limit model complexity, or complicate debugging and validation. Furthermore, security concerns extend to adversarial threats, including data poisoning and model inversion attacks, which must be actively mitigated in clinical deployments.

8.5 Infrastructure and Resource Constraints

The feasibility of deploying ophthalmic AI varies significantly across healthcare settings. While high‐resource environments may support advanced computational infrastructure and seamless connectivity, many regions—particularly in low‐ and middle‐income countries—face substantial limitations. These constraints include unreliable internet access, variability in imaging equipment, and limited technical support. In such contexts, cloud‐dependent solutions may be impractical. Instead, there is growing interest in lightweight, on‐device inference systems that can operate independently of continuous network connectivity. Edge‐based approaches, supported by model compression and efficient architectures, enable real‐time analysis at the point of care. This is particularly relevant for large‐scale screening initiatives, where rapid decision‐making is required in resource‐constrained environments. Nevertheless, maintaining consistent performance across varying image quality and acquisition conditions remains a key technical challenge.

8.6 Economic and Societal Impact

Beyond technical feasibility, the long‐term value of ophthalmic AI must be evaluated in terms of its economic and societal implications. From a health systems perspective, AI has the potential to reduce diagnostic workload, streamline screening programs, and enable earlier detection of vision‐threatening conditions. These benefits, however, must be balanced against the costs associated with development, deployment, maintenance, and training. Cost‐effectiveness is often context‐dependent. In high‐prevalence settings, automated screening may yield significant savings by reducing specialist burden. In contrast, in low‐prevalence populations, excessive false positives may offset these gains. Therefore, economic evaluation should consider not only direct costs but also downstream effects on healthcare utilization and patient outcomes. At a societal level, the deployment of AI raises broader questions regarding equity and access. While AI has the potential to democratize eye care, there is also a risk that technological disparities may widen existing gaps if advanced systems remain concentrated in well‐resourced settings. Ethical considerations, including transparency, accountability, and the role of clinicians in decision‐making, must therefore remain central to deployment strategies. A structured summary of key translational challenges, their underlying factors, and clinical implications is provided in Table 7. The table emphasizes that successful deployment of ophthalmic AI depends on coordinated optimization across technical, clinical, regulatory, and socioeconomic dimensions rather than algorithmic performance alone.

9 Future Directions

The field of AI in ophthalmology is undergoing a rapid transition from task‐specific, data‐intensive models toward more generalized, adaptive, and clinically integrated systems. Emerging paradigms are increasingly shaped by advances in foundation models, multimodal reasoning, edge deployment, and personalized healthcare. These developments signal a shift from isolated algorithmic performance toward holistic, patient‐centered intelligence embedded within real world clinical ecosystems.

9.1 Foundation Models and Generative AI

The emergence of foundation models marks a significant shift in the trajectory of AI in ophthalmology. Unlike earlier task‐specific systems, these models are trained on large and diverse datasets, enabling them to capture generalizable visual and semantic representations. As a result, they can be adapted to a wide range of downstream clinical tasks with minimal additional supervision. In practical terms, this paradigm offers several compelling advantages. Foundation models can support few‐shot or even zero‐shot learning, which is particularly valuable for rare ophthalmic conditions where labeled data are scarce. They also demonstrate improved robustness across variations in imaging devices, acquisition protocols, and patient populations—an essential requirement for real‐world deployment. Moreover, the integration of vision and language capabilities opens the possibility of automated report generation, clinical summarization, and decision support that more closely resembles human reasoning. Generative AI further extends these capabilities by enabling synthesis rather than mere prediction. For example, generative models can assist in creating structured clinical narratives from imaging data, simulate plausible disease progression scenarios, or augment datasets in situations where annotated data are limited. However, these advances are accompanied by new risks. Hallucinated outputs, lack of grounding in clinical evidence, and overgeneralization remain important concerns. Consequently, future research must emphasize alignment with validated medical knowledge, rigorous evaluation, and safeguards against unsafe or misleading outputs.

9.2 Multimodal Clinical Intelligence

Clinical decision‐making in ophthalmology rarely relies on a single source of information. Instead, clinicians integrate imaging findings with patient history, symptoms, functional assessments, and, increasingly, genetic information. Current AI systems, however, often operate within narrow modality specific silos, limiting their ability to replicate this holistic reasoning process. Multimodal AI seeks to bridge this gap by learning unified representations from heterogeneous data sources, including fundus images, OCT scans, VF tests, EHRs, and genomic profiles. Such integration has the potential to significantly enhance diagnostic accuracy and enable more nuanced disease characterization. For example, combining structural information from OCT with functional deficits observed in VFs can improve the assessment of glaucoma progression. Similarly, incorporating demographic and genetic risk factors may lead to more accurate prediction models for conditions such as AMD. Despite these advantages, multimodal learning introduces several technical and practical challenges. Data from different modalities often vary in scale, resolution, and availability, and missing data are common in clinical interpretability becomes increasingly difficult. Addressing these challenges will be essential to ensure that multimodal systems remain both reliable and clinically interpretable.

9.3 Edge AI and Point‐of‐Care Systems

One of the most pressing challenges in global ophthalmology is the unequal distribution of healthcare resources. Specialist services are often concentrated in urban centers, leaving rural and underserved populations with limited access to timely diagnosis and treatment. Edge AI offers a promising solution by enabling real‐time inference directly on portable devices deployed at the point of care. Rather than relying on centralized cloud infrastructure, edge‐based systems perform computations locally using compact and efficient models. This approach reduces latency, minimizes dependence on high bandwidth connectivity, and enhances data privacy. In practical scenarios, such systems can support community‐level screening programs, mobile eye clinics, and teleophthalmology initiatives. Recent advances in model compression, quantization, and hardware‐aware design have made it increasingly feasible to deploy sophisticated AI models on resource‐constrained devices. Nevertheless, ensuring consistent performance under varying environmental conditions—such as changes in lighting, image quality, or device calibration—remains a critical challenge. In this context, robustness and reliability are as important as efficiency. From a broader perspective, edge AI represents not just a technological innovation but a pathway toward more equitable healthcare delivery.

9.4 Personalized Ophthalmology

The paradigm of ophthalmic care is gradually shifting from reactive treatment to proactive and personalized management. AI plays a central role in this transition by enabling predictive modeling of disease trajectories and treatment outcomes. Rather than applying uniform treatment protocols, personalized ophthalmology aims to tailor interventions based on individual patient characteristics, including imaging biomarkers, clinical history, and temporal disease patterns. For instance, AI models can be used to predict response to anti‐VEGF therapy in AMD or to estimate the rate of glaucoma progression based on longitudinal data. Such predictive capabilities allow clinicians to optimize follow‐up intervals, prioritize high‐risk patients, and reduce unnecessary interventions. At the same time, they introduce new complexities. Personalized models require high‐quality longitudinal datasets, careful handling of temporal variability, and mechanisms to ensure stability over time. Ultimately, the goal is to move toward a form of precision ophthalmology in which clinical decisions are informed by data‐driven insights tailored to each patient.

9.5 Prospective Trials and Real‐World Evidence

While many ophthalmic AI systems have demonstrated impressive performance on curated datasets, their real‐world effectiveness remains insufficiently validated. A critical limitation of current research is the reliance on retrospective studies conducted under controlled conditions, which may not reflect the variability encountered in clinical practice. Bridging this gap requires a shift toward prospective evaluation and real‐world evidence generation. Multi‐center studies involving diverse populations and imaging devices are essential to assess generalizability. Longitudinal trials can provide insights into how AI systems perform over time and under changing clinical conditions. Equally important is the evaluation of AI systems within actual clinical workflows. This includes understanding how clinicians interact with AI outputs, how decisions are influenced, and whether patient outcomes improve as a result. Metrics such as diagnostic accuracy must therefore be complemented by measures of usability, trust, efficiency, and cost effectiveness. A structured overview of emerging paradigms, their capabilities, clinical value, and associated challenges is summarized in Table 8. The table highlights that future progress in ophthalmic AI depends on balancing innovation with robustness, interpretability, and real‐world validation.

10 Limitations of This Review

This review has several limitations that should be considered when interpreting its scope and conclusions. First, it was designed as a structured narrative review rather than a full systematic review or meta‐analysis. Therefore, although the literature search and synthesis were organized around predefined themes, databases, keywords, and inclusion priorities, the review did not perform formal quantitative pooling, risk of‐bias scoring, or meta‐analytic comparison of diagnostic performance across studies.

Second, the search and prioritization strategy was intentionally oriented toward breadth, clinical relevance, and translational significance. Priority was given to studies that were methodologically influential, clinically meaningful, externally validated, prospectively evaluated, or relevant to real‐world deployment. As a result, the review may not include every technical variant, algorithmic refinement, or narrowly focused experimental study in ophthalmic AI. This choice was made to maintain emphasis on the relationship among AI methods, ophthalmic use cases, evidence maturity, and clinical implementation.

Third, the literature on foundation models, generative AI, multimodal learning, and vision‐language systems is evolving rapidly. New models, datasets, benchmarks, and regulatory discussions may emerge quickly, and some conclusions regarding these paradigms may require revision as stronger clinical evidence becomes available. In particular, the current evidence base for generative and foundation‐model‐based ophthalmic AI remains less mature than that for conventional image‐based screening applications.

Fourth, heterogeneity across datasets, imaging devices, annotation protocols, disease definitions, validation strategies, and reporting standards limited direct quantitative comparison among studies. Reported values for accuracy, sensitivity, specificity, AUC, Dice score, calibration, and other metrics are often not directly comparable because they may be derived from different populations, disease prevalences, image‐quality criteria, and decision thresholds. For this reason, performance estimates were interpreted qualitatively and contextually rather than pooled statistically.

Finally, the maturity of evidence differs substantially across ophthalmic domains. DR screening has relatively stronger validation and deployment evidence, whereas areas such as anterior segment AI, infectious keratitis, rare ocular diseases, multimodal treatment guidance, and generative clinical reporting remain less mature. Consequently, conclusions regarding these emerging areas should be viewed as cautious and provisional, pending larger external validations, prospective trials, and real‐world implementation studies.

11 Conclusion

AI has become a major force in ophthalmology because the specialty is strongly supported by imaging, structured measurements, and repeatable diagnostic workflows. This review has synthesized the field across methodological foundations, ophthalmic imaging modalities, public datasets, disease‐specific applications, evaluation practices, translational barriers, and emerging directions. The current evidence shows that AI has achieved its greatest maturity in image‐based screening tasks, particularly DR, while important progress is also evident in glaucoma assessment, AMD monitoring, cataract grading, infectious keratitis recognition, keratoconus detection, segmentation, and multimodal prediction. Nevertheless, high retrospective performance should not be equated with clinical readiness. Many ophthalmic AI systems remain limited by narrow datasets, insufficient external validation, device specific performance variation, incomplete subgroup analysis, uncertain calibration, and limited evidence from prospective or real‐world studies. Therefore, the central question is no longer whether AI can achieve high accuracy under controlled conditions, but whether it can remain safe, reliable, explainable, and clinically useful across diverse patients, imaging devices, healthcare settings, and workflows. The next phase of ophthalmic AI should move beyond benchmark‐oriented reporting toward clinically grounded validation. Future studies should prioritize external and multicenter testing, prospective trials, calibration assessment, uncertainty‐aware decision support, bias analysis, workflow integration, and post‐deployment monitoring. Thresholds should be selected according to clinical context, balancing missed disease against false referral burden. Similarly, AI outputs should support, rather than replace, clinician judgment, especially in high‐risk or ambiguous cases. Emerging paradigms such as foundation models, generative AI, multimodal learning, and edge deployment may expand the reach and functionality of ophthalmic AI, particularly in underserved and resource‐constrained settings. However, these systems must be developed with caution, given risks related to hallucination, weak clinical grounding, accountability, privacy, and performance drift. Their value will depend not merely on technical sophistication, but on demonstrable improvements in patient outcomes, access to care, and clinical efficiency.

References

[1]

Z. Li, L. Wang, X. Wu, et al., “Artificial Intelligence in Ophthalmology: The Path to the Real‐World Clinic,” Cell Reports Medicine 4, no. 7 (2023): 101095.

[2]

J. Wilson, Leading the AI Ophthalmology Revolution [Internet] (Johns Hopkins Medicine News & Stories; 2024).

[3]

H. Maehara, Y. Ueno, T. Yamaguchi, et al., “Artificial Intelligence Support Improves Diagnosis Accuracy in Anterior Segment Eye Diseases,” Scientific Reports 15, no. 1 (2025): 5117.

[4]

M. Miyake, M. Akiyama, K. Kashiwagi, T. Sakamoto, and T. Oshika, “Japan Ocular Imaging Registry: A National Ophthalmology Real‐World Database,” Japanese Journal of Ophthalmology 66, no. 6 (2022): 499–503.

[5]

H. Tabuchi, J. Engelmann, F. Maeda, et al., “Using Artificial Intelligence to Improve Human Performance: Efficient Retinal Disease Detection Training With Synthetic Images,” British Journal of Ophthalmology 108, no. 10 (2024): 1430–1435.

[6]

X. Chen, R. Chen, P. Xu, et al., “From Visual Question Answering to Intelligent AI Agents in Ophthalmology,” British Journal of Ophthalmology 110, no. 1 (2026): 1–7.

[7]

D. Vamsidhar, S. Kolhar, S. Patil, and S. Kumar, “Advancements in Ophthalmology Healthcare Using Multimodal AI: A Systematic Review of Methods, Applications, and Future Directions,” Discover Artificial Intelligence 6, no. 1 (2026): 236.

[8]

M. Siwik, N. Nalecz, Z. Jankowska, A. Laudencka, K. Kazmierczak, and B. Kaluzny, “Application of Artificial Intelligence in Glaucoma Diagnosis: A Literature Review,” Okulistyka 28, no. 3 (2026): 49–53.

[9]

B. T. Soetikno, C. S. Nielsen, A. Pollreisz, and D. S. Ting, “Toward Autonomous Discovery: Agentic AI and the Future of Ophthalmic Research,” Current Opinion in Ophthalmology 37, no. 1 (2026): 60–65.

[10]

R. H. Barcelo‐Canton, M. Yu, C. Liu, A. Takahashi, I. X. Lee, and Y. C. Liu, “Applications of Artificial Intelligence in Corneal Nerve Images in Ophthalmology,” Diagnostics 16, no. 4 (2026): 602.

[11]

E. Garcia, M. K. Adam, E. Dow, et al., “Artificial Intelligence in Clinical Care: Perceptions of Retina Specialists and Patients Gathered Through a Multicenter Survey,” Journal of VitreoRetinal Diseases 10, no. 1 (2026): 68–73.

[12]

F. Buckmaster, D. van Staden, and L. Coetzee, “Attitudes and Knowledge Levels of Optometry Students and Educators Towards Artificial Intelligence in Optometric Practice: An Online Cross‐Sectional Survey,” Ophthalmic and Physiological Optics 46, no. 2 (2026): 419–428.

[13]

K. Gharbi, P. van Wijngaarden, and X. Hadoux, “Foundation Models for Ophthalmic Imaging,” Survey of Ophthalmology 71, no. 4 (2026): 1129–1147.

[14]

A. Pratap, Y. T. Huang, C. K. Chiang, et al., “Deep Learning Technology in Genomics, Radiotherapy, and Ophthalmology for Precision Medicine,” Journal of Physiological Investigation 69, no. 2 (2026): 127–156.

[15]

W. Tang, Y. Hang, J. Zhang, and Y. Dang, “Recent Advances and Future Directions of Artificial Intelligence in Glaucoma Management,” Open Ophthalmology Journal 20, no. 1 (2026): e18743641430437.

[16]

A. Grzybowski, K. Zhao, and K. Jin, “Agentic Artificial Intelligence in Ophthalmology: Toward Autonomous, Adaptive, and Ethical Eye Care,” Acta Ophthalmologica (2026): 1–6.

[17]

M. N. Radeva, E. Hristova, R. T. Georgiev, and Z. I. Zlatarova, “Awareness, Trust, and Expectations of AI for Glaucoma Care Among Bulgarian Ophthalmologists: Role of Demographic Factors,” PLOS Digital Health 5, no. 1 (2026): e0001199.

[18]

K. Jin, T. Yu, and A. Grzybowski, “Multimodal Artificial Intelligence in Ophthalmology: Applications, Challenges, and Future Directions,” Survey of Ophthalmology 71, no. 1 (2025): 158–167.

[19]

A. R. Ran, C. H. Lui, Y. C. Tham, et al., “The Acceptance of Ophthalmic Artificial Intelligence for Eye Diseases: A Literature Review and Qualitative Analysis,” Eye 39, no. 12 (2025): 2353–2362.

[20]

S. Chen and W. Bai, “Artificial Intelligence Technology in Ophthalmology Public Health: Current Applications and Future Directions,” Frontiers in Cell and Developmental Biology 13 (2025): 1576465.

[21]

P. Heydon, C. Egan, L. Bolter, et al., “Prospective Evaluation of an Artificial Intelligence‐Enabled Algorithm for Automated Diabetic Retinopathy Screening of 30 000 Patients,” British Journal of Ophthalmology 105, no. 5 (2021): 723–728.

[22]

V. Gulshan, R. P. Rajan, K. Widner, et al., “Performance of a Deep‐Learning Algorithm vs Manual Grading for Detecting Diabetic Retinopathy in India,” JAMA Ophthalmology 137, no. 9 (2019): 987–993.

[23]

R. Fan, C. Bowd, M. Christopher, et al., “Detecting Glaucoma in the Ocular Hypertension Study Using Deep Learning,” JAMA Ophthalmology 140, no. 4 (2022): 383–391.

[24]

F. A. Medeiros, A. A. Jammal, and E. B. Mariottoni, “Detection of Progressive Glaucomatous Optic Nerve Damage on Fundus Photographs With Deep Learning,” Ophthalmology 128, no. 3 (2021): 383–392.

[25]

R. Asaoka, H. Murata, A. Iwase, and M. Araie, “Detecting Preperimetric Glaucoma With Standard Automated Perimetry Using a Deep Learning Classifier,” Ophthalmology 123, no. 9 (2016): 1974–1980.

[26]

Q. Lu, L. Wei, W. He, et al., “Lens Opacities Classification System III‐Based Artificial Intelligence Program for Automatic Cataract Grading,” Journal of Cataract & Refractive Surgery 48, no. 5 (2022): 528–534.

[27]

M. Tiwari, C. Piech, M. Baitemirova, et al., “Differentiation of Active Corneal Infections From Healed Scars Using Deep Learning,” Ophthalmology 129, no. 2 (2022): 139–146.

[28]

P. Zeboulon, G. Debellemanière, M. Bouvet, and D. Gatinel, “Corneal Topography Raw Data Classification Using a Convolutional Neural Network,” American Journal of Ophthalmology 219, no. 1 (2020): 33–39.

[29]

P. M. Burlina, N. Joshi, M. Pekala, K. D. Pacheco, D. E. Freund, and N. M. Bressler, “Automated Grading of Age‐Related Macular Degeneration From Color Fundus Images Using Deep Convolutional Neural Networks,” JAMA Ophthalmology 135, no. 11 (2017): 1170–1176.

[30]

H. Zhang, K. Niu, Y. Xiong, W. Yang, Z. He, and H. Song, “Automatic Cataract Grading Methods Based on Deep Learning,” Computer Methods and Programs in Biomedicine 182 (2019): 104978.

[31]

O. Srivastava, M. Tennant, P. Grewal, U. Rubin, and M. Seamone, “Artificial Intelligence and Machine Learning in Ophthalmology: A Review,” Indian Journal of Ophthalmology 71, no. 1 (2023): 11–17.

[32]

K. Jin and J. Ye, “Artificial Intelligence and Deep Learning in Ophthalmology: Current Status and Future Perspectives,” Advances in Ophthalmology Practice and Research 2, no. 3 (2022): 100078.

[33]

J. H. Wu, N. D. Koseoglu, C. Jones, and T. A. Liu, “Vision Transformers: The Next Frontier for Deep Learning‐Based Ophthalmic Image Analysis,” Saudi Journal of Ophthalmology 37, no. 3 (2023): 173–178.

[34]

H. Rocha, Y. J. Chong, A. J. Thirunavukarasu, et al., “Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmo‐Logical Questions,” JAMA Ophthalmology 144, no. 1 (2026): 5–13.

[35]

M. A. Chia, F. Antaki, Y. Zhou, A. W. Turner, A. Y. Lee, and P. A. Keane, “Foundation Models in Ophthalmology,” British Journal of Ophthalmology 108, no. 10 (2024): 1341–1348.

[36]

H. Shah, R. Patel, S. Hegde, and H. Dalvi, “XAI Meets Ophthalmology: An Explainable Approach to Cataract Detection Using VGG‐19 and Grad‐CAM,” in 2023 IEEE Pune Section International Conference (PuneCon) (IEEE, 2023), 1–8.

[37]

E. Dugas, J. Jared, and W. Cukierski, Diabetic Retinopathy Detection (Kaggle, 2015).

[38]

M. Karthik and S. Dane, APTOS 2019 Blindness Detection (Kaggle, 2019).

[39]

E. Decencière, X. Zhang, G. Cazuguel, et al., “Feedback on a Publicly Distributed Database: The Messidor Database,” Image Analysis & Stereology 33, no. 3 (2014): 231–234.

[40]

M. D. Abràmoff, J. C. Folk, D. P. Han, et al., “Automated Analysis of Retinal Images for Detection of Referable Diabetic Retinopathy,” JAMA Ophthalmology 131, no. 3 (2013): 351–357.

[41]

P. Porwal, S. Pachade, R. Kamble, et al., “Indian Diabetic Retinopathy Image Dataset (IDRiD): A Database for Diabetic Retinopathy Screening Research,” Data 3, no. 3 (2018): 25.

[42]

Grand Challenge, ODIR‐2019: Ocular Disease Intelligent Recognition [Internet] (Dataset/challenge page).

[43]

H. Fang, F. Li, H. Fu, et al., ADAM Challenge: Detecting age‐related Macular Degeneration From Fundus Images [Internet] (Challenge resource).

[44]

Grand Challenge, PALM: Pathologic Myopia Challenge [Internet] (Dataset page).

[45]

J. I. Orlando, H. Fu, J. Barbosa Breda, et al., “REFUGE Challenge: A Unified Framework for Evaluating Automated Methods for Glaucoma Assessment From Fundus Photographs,” Medical Image Analysis 59 (2020): 101570.

[46]

Z. Zhang, F. S. Yin, and J. Liu, et al., “ORIGA‐light: An Online Retinal Fundus Image Database for Glaucoma Analysis and Research,” in Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) (Institute of Electrical and Electronics Engineers (IEEE), 2010): 3065–3068.

[47]

A. Almazroa, S. Alodhayb, E. Osman, et al., Retinal Fundus Images for Glaucoma Analysis: The RIGA Dataset (Dataset record, University of Michigan Deep Blue, 2018).

[48]

J. Wu, H. Fang, F. Li, et al., “GAMMA Challenge: Glaucoma Grading From Multi‐Modality Images,” Medical Image Analysis 88 (2023): 102938.

[49]

H. Bogunović, F. Venhuizen, S. Klimscha, et al., “RETOUCH: The Retinal OCT Fluid Detection and Segmentation Benchmark and Challenge,” IEEE Transactions on Medical Imaging 38, no. 8 (2019): 1858–1874.

[50]

S. J. Chiu, M. J. Allingham, P. S. Mettu, et al., “Kernel Regression Based Segmentation of Optical Coherence Tomography Images With Diabetic Macular Edema,” Biomedical Optics Express 6, no. 4 (2015): 1172–1194.

[51]

J. J. Staal, M. D. Abramoff, M. Niemeijer, M. A. Viergever, and B. van Ginneken, “Ridge Based Vessel Segmentation in Color Images of the Retina,” IEEE Transactions on Medical Imaging 23, no. 4 (2004): 501–509.

[52]

A. Hoover, V. Kouznetsova and M. Goldbaum, “Locating Blood Vessels in Retinal Images by Piece‐Wise Threshold Probing of a Matched Filter Response,” IEEE Transactions on Medical Imaging 19, no. 3 (2000): 203–210.

[53]

S. Barman, A. Hoppe, P. Remagnino, et al., CHASE_DB1 Retinal Vessel Reference Dataset (Kingston University, 2012), CHASEDB1(.zip), readme(.txt).

[54]

Q. Hu, M. D. Abràmoff, and M. K. Garvin, “Automated Separation of Binary Overlapping Trees in Low‐Contrast Color Retinal Images,” Medical Image Computing and Computer‐Assisted Intervention 16, no. 2 (2013): 436–443.

[55]

Y. Ma, H. Hao, J. Xie, et al., “ROSE: A Retinal OCT‐Angiography Vessel Segmentation Dataset and New Model,” IEEE Transactions on Medical Imaging 40, no. 3 (2021): 928–939.

[56]

A. Grzybowski, K. Jin, J. Zhou, et al., “Retina Fundus Photograph‐Based Artificial Intelligence Algorithms in Medicine: A Systematic Review,” Ophthalmology and Therapy 13, no. 8 (2024): 2125–2149.

[57]

D. Restrepo, J. M. Quion, F. Do Carmo Novaes, et al., “Ophthalmology Optical Coherence Tomography Databases for Artificial Intelligence Algorithm: A Review,” Seminars in Ophthalmology 39, no. 3 (2024): 193–200.

[58]

T. Murata, T. Hirano, H. Mizobe, and S. Toba, “OCT‐Angiography Based Artificial Intelligence‐Inferred Fluorescein Angiography for Leakage Detection in Retina,” Biomedical Optics Express 14, no. 11 (2023): 5851–5860.

[59]

E. Shimizu, K. Tanaka, H. Nishimura, et al., “The Use of Artificial Intelligence for Estimating Anterior Chamber Depth From Slit‐Lamp Images Developed Using Anterior‐Segment Optical Coherence Tomography,” Bioengineering 11, no. 10 (2024): 1005.

[60]

Z. Zhou, B. Li, J. Su, et al., “An Artificial Intelligence Model for the Simulation of Visual Effects in Patients With Visual Field Defects,” Annals of Translational Medicine 8, no. 11 (2020): 703.

[61]

M. Jiménez‐García, I. Issarti, E. O. Kreps, et al., “Forecasting Progressive Trends in Keratoconus by Means of a Time Delay Neural Network,” Journal of Clinical Medicine 10, no. 15 (2021): 3238.

[62]

D. S. W. Ting, C. Y. L. Cheung, G. Lim, et al., “Development and Validation of a Deep Learning System for Diabetic Retinopathy and Related Eye Diseases Using Retinal Images From Multiethnic Populations With Diabetes,” JAMA 318, no. 22 (2017): 2211–2223.

[63]

V. Gulshan, L. Peng, M. Coram, et al., “Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs,” JAMA 316, no. 22 (2016): 2402–2410.

[64]

Z. Li, S. Keel, C. Liu, et al., “An Automated Grading System for Detection of Vision‐Threatening Referable Diabetic Retinopathy on the Basis of Color Fundus Photographs,” Diabetes Care 41, no. 12 (2018): 2509–2516.

[65]

X. Liu, T. K. Ali, P. Singh, et al., “Deep Learning to Detect OCT‐Derived Diabetic Macular Edema From Color Retinal Photographs: A Multicenter Validation Study,” Ophthalmology Retina 6, no. 5 (2022): 398–410.

[66]

T. Araujo, G. Aresta, L. Mendonça, et al., “DR|GRADUATE: Uncertainty‐Aware Deep Learning‐Based Diabetic Retinopathy Grading in Eye Fundus Images,” Medical Image Analysis 63 (2020): 101715.

[67]

L. Dai, L. Wu, H. Li, et al., “A Deep Learning System for Detecting Diabetic Retinopathy Across the Disease Spectrum,” Nature Communications 12, no. 1 (2021): 3242.

[68]

A. Y. Lee, R. T. Yanagihara, C. S. Lee, et al., “Multicenter, Head‐to‐Head, Real‐World Validation Study of Seven Automated Artificial Intelligence Diabetic Retinopathy Screening Systems,” Diabetes Care 44, no. 5 (2021): 1168–1175.

[69]

S. Natarajan, A. Jain, R. Krishnan, A. Rogye, and S. Sivaprasad, “Diagnostic Accuracy of Community‐Based Diabetic Retinopathy Screening With an Offline Artificial Intelligence System on a Smartphone,” JAMA Ophthalmology 137, no. 10 (2019): 1182–1188.

[70]

J. Xiong, F. Li, D. Song, et al., “Multimodal Machine Learning Using Visual Fields and Peripapillary Circular OCT Scans in Detection of Glaucomatous Optic Neuropathy,” Ophthalmology 129, no. 2 (2022): 171–180.

[71]

A. Dixit, J. Yohannan, and M. V. Boland, “Assessing Glaucoma Progression Using Machine Learning Trained on Longitudinal Visual Field and Clinical Data,” Ophthalmology 128, no. 7 (2021): 1016–1026.

[72]

F. Li, Y. Su, F. Lin, et al., “A Deep‐Learning System Predicts Glaucoma Incidence and Progression Using Retinal Photographs,” Journal of Clinical Investigation 132, no. 11 (2022): e157968.

[73]

A. R. Ran, C. Y. Cheung, X. Wang, et al., “Detection of Glaucomatous Optic Neuropathy With Spectral‐Domain Optical Coherence Tomography: A Retrospective Training and Validation Deep Learning Analysis,” Lancet Digital Health 1, no. 4 (2019): e172–e182.

[74]

Z. Li, C. Guo, D. Lin, et al., “Deep Learning for Automated Glaucomatous Optic Neuropathy Detection From Ultra‐Widefield Fundus Images,” British Journal of Ophthalmology 105, no. 11 (2021): 1548–1554.

[75]

F. Li, Y. Yang, X. Sun, et al., “Digital Gonioscopy Based on Three‐Dimensional Anterior‐Segment OCT: An International Multicenter Study,” Ophthalmology 129, no. 1 (2022): 45–53.

[76]

S. Yousefi, T. Elze, L. R. Pasquale, et al., “Monitoring Glaucomatous Functional Loss Using an Artificial Intelligence‐Enabled Dashboard,” Ophthalmology 127, no. 9 (2020): 1170–1178.

[77]

K. R. Martin, K. Mansouri, R. N. Weinreb, et al., “Use of Machine Learning on Contact Lens Sensor‐Derived Parameters for the Diagnosis of Primary Open‐Angle Glaucoma,” American Journal of Ophthalmology 194 (2018): 46–53.

[78]

D. K. Hwang, C. C. Hsu, K. J. Chang, et al., “Artificial Intelligence‐Based Decision‐Making For Age‐Related Macular Degeneration,” Theranostics 9, no. 1 (2019): 232–245.

[79]

D. S. Kermany, M. Goldbaum, W. Cai, et al., “Identifying Medical Diagnoses and Treatable Diseases by Image‐Based Deep Learning,” Cell 172, no. 5 (2018): 1122–1131.e9.

[80]

Y. Peng, S. Dharssi, Q. Chen, et al., “DeepSeeNet: A Deep Learning Model for Automated Classification of Patient‐Based Age‐Related Macular Degeneration Severity From Color Fundus Photographs,” Ophthalmology 126, no. 4 (2019): 565–575.

[81]

S. Keel, Z. Li, J. Scheetz, et al., “Development and Validation of a Deep‐Learning Algorithm for the Detection of Neovascular Age‐Related Macular Degeneration From Colour Fundus Photographs,” Clinical and Experimental Ophthalmology 47, no. 8 (2019): 1009–1018.

[82]

F. Grassmann, J. Mengelkamp, C. Brandl, et al., “A Deep Learning Algorithm for Prediction of Age‐Related Eye Disease Study Severity Scale for Age‐Related Macular Degeneration From Color Fundus Photography,” Ophthalmology 125, no. 9 (2018): 1410–1420.

[83]

N. Rakocz, J. N. Chiang, M. G. Nittala, et al., “Automated Identification of Clinical Features From Sparsely Annotated 3‐Dimensional Medical Imaging,” npj Digital Medicine 4, no. 1 (2021): 44.

[84]

J. Yim, R. Chopra, T. Spitz, et al., “Predicting Conversion to Wet Age‐Related Macular Degeneration Using Deep Learning,” Nature Medicine 26, no. 6 (2020): 892–899.

[85]

I. Potapenko, B. Thiesson, M. Kristensen, et al., “Automated Artificial Intelligence‐Based System for Clinical Follow‐Up of Patients With Age‐Related Macular Degeneration,” Acta Ophthalmologica 100, no. 8 (2022): 927–936.

[86]

B. Yellapragada, S. Hornauer, K. Snyder, S. Yu, and G. Yiu, “Self‐Supervised Feature Learning and Phenotyping for Assessing Age‐Related Macular Degeneration Using Retinal Fundus Images,” Ophthalmology Retina 6, no. 2 (2022): 116–129.

[87]

X. Wu, Y. Huang, Z. Liu, et al., “Universal Artificial Intelligence Platform for Collaborative Management of Cataracts,” British Journal of Ophthalmology 103, no. 11 (2019): 1553–1560.

[88]

X. Xu, J. Li, Y. Guan, et al., “GLA‐Net: A Global‐Local Attention Network for Automatic Cataract Classification,” Journal of Biomedical Informatics 124 (2021): 103939.

[89]

X. Xu, L. Zhang, J. Li, Y. Guan, and L. Zhang, “A Hybrid Global‐Local Representation CNN Model for Automatic Cataract Grading,” IEEE Journal of Biomedical and Health Informatics 24, no. 2 (2020): 556–567.

[90]

Y. C. Tham, J. H. L. Goh, A. Anees, et al., “Detecting Visually Significant Cataract Using Retinal Photograph‐Based Deep Learning,” Nature Aging 2, no. 3 (2022): 264–271.

[91]

H. Lin, R. Li, Z. Liu, et al., “Diagnostic Efficacy and Therapeutic Decision‐Making Capacity of an Artificial Intelligence Platform for Childhood Cataracts in Eye Clinics: A Multicentre Randomized Controlled Trial,” EClinicalMedicine 9 (2019): 52–59.

[92]

D. Lin, J. Chen, Z. Lin, et al., “A Practical Model for the Identification of Congenital Cataracts Using Machine Learning,” EBioMedicine 51 (2020): 102621.

[93]

T. D. L. Keenan, Q. Chen, E. Agrón, et al., “DeepLensNet: Deep Learning Automated Diagnosis and Quantitative Classification of Cataract Type and Severity,” Ophthalmology 129, no. 5 (2022): 571–584.

[94]

X. Gao, S. Lin, and T. Y. Wong, “Automatic Feature Learning to Grade Nuclear Cataracts Based on Deep Learning,” IEEE Transactions on Biomedical Engineering 62, no. 11 (2015): 2693–2701.

[95]

H. Gu, Y. Guo, L. Gu, et al., “Deep Learning for Identifying Corneal Diseases From Ocular Surface Slit‐Lamp Photographs,” Scientific Reports 10, no. 1 (2020): 17851.

[96]

Z. Li, J. Jiang, K. Chen, et al., “Preventing Corneal Blindness Caused by Keratitis Using Artificial Intelligence,” Nature Communications 12, no. 1 (2021): 3738.

[97]

A. K. Ghosh, R. Thammasudjarit, P. Jongkhajornpong, J. Attia, and A. Thakkinstian, “Deep Learning for Discrimination Between Fungal Keratitis and Bacterial Keratitis: Deepkeratitis,” Cornea 41, no. 5 (2022): 616–622.

[98]

T. K. Redd, N. V. Prajna, M. Srinivasan, et al., “Image‐Based Differentiation of Bacterial and Fungal Keratitis Using Deep Convolutional Neural Networks,” Ophthalmology Science 2 (2022): 100119.

[99]

Y. Xu, M. Kong, W. Xie, et al., “Deep Sequential Feature Learning in Clinical Image Classification of Infectious Keratitis,” Engineering 7 (2021): 1002–1010.

[100]

L. Wang, K. Chen, H. Wen, et al., “Feasibility Assessment of Infectious Keratitis Depicted on Slit‐Lamp and Smartphone Photographs Using Deep Learning,” International Journal of Medical Informatics 155 (2021): 104583.

[101]

J. Lv, K. Zhang, Q. Chen, et al., “Deep Learning‐Based Automated Diagnosis of Fungal Keratitis With In Vivo Confocal Microscopy Images,” Annals of Translational Medicine 8, no. 11 (2020): 706.

[102]

W. Wu, S. Huang, X. Xie, et al., “Raman Spectroscopy May Allow Rapid Noninvasive Screening of Keratitis and Conjunctivitis,” Photodiagnosis and Photodynamic Therapy 37 (2022): 102689.

[103]

Z. Ren, W. Li, Q. Liu, Y. Dong, and Y. Huang, “Profiling of the Conjunctival Bacterial Microbiota Reveals the Feasibility of Utilizing a Microbiome‐Based Machine Learning Model to Differentially Diagnose Microbial Keratitis and the Core Components of the Conjunctival Bacterial Interaction Network,” Frontiers in Cellular and Infection Microbiology 12 (2022): 860370.

[104]

Y. Xie, L. Zhao, X. Yang, et al., “Screening Candidates for Refractive Surgery With Corneal Tomographic‐Based Deep Learning,” JAMA Ophthalmology 138, no. 5 (2020): 519–526.

[105]

A. H. Al‐Timemy, Z. M. Mosa, Z. Alyasseri, et al., “A Hybrid Deep Learning Construct for Detecting Keratoconus From Corneal Maps,” Translational Vision Science & Technology 10, no. 14 (2021): 16.

[106]

K. Cao, K. Verspoor, E. Chan, M. Daniell, S. Sahebjada, and P. N. Baird, “Machine Learning With a Reduced Dimensionality Representation of Comprehensive Pentacam Tomography Parameters to Identify Subclinical Keratoconus,” Computers in Biology and Medicine 138 (2021): 104884.

[107]

G. Castro‐Luna, D. Jiménez‐Rodríguez, A. B. Castaño‐Fernández, and A. Perez‐Rueda, “Diagnosis of Subclinical Keratoconus Based on Machine Learning Techniques,” Journal of Clinical Medicine 10, no. 18 (2021): 4281.

[108]

I. Ruiz Hidalgo, P. Rodriguez, J. J. Rozema, et al., “Evaluation of a Machine Learning Classifier for Keratoconus Detection Based on Scheimpflug Tomography,” Cornea 35, no. 6 (2016): 827–832.

[109]

C. Shi, M. Wang, T. Zhu, et al., “Machine Learning Helps Improve Diagnostic Ability of Subclinical Keratoconus Using Scheimpflug and OCT Imaging Modalities,” Eye and Vision 7, no. 1 (2020): 48.

[110]

G. C. Almeida, R. C. Guido, H. M. B. Silva, et al., “New Artificial Intelligence Index Based on Scheimpflug Corneal Tomography to Distinguish Subclinical Keratoconus From Healthy Corneas,” Journal of Cataract & Refractive Surgery 48, no. 10 (2022): 1168–1174.

[111]

I. Issarti, A. Consejo, M. Jiménez‐García, S. Hershko, C. Koppen, and J. J. Rozema, “Computer Aided Diagnosis for Suspect Keratoconus Detection,” Computers in Biology and Medicine 109 (2019): 33–42.

Rights & permissions

2026 The Author(s). Eye & ENT Research published by John Wiley & Sons Australia, Ltd on behalf of Higher Education Press.

PDF (1190KB)

13

Accesses

0

Citation

Detail

Sections
Recommended

/