Introduction
Ocular trauma is a leading cause of blinding eye diseases worldwide. When the globe is subjected to blunt impact, a tearing injury occurs at the junction between the iris root and the ciliary body, leading to a pathological posterior displacement of the anterior chamber angle structures—a condition defined as angle recession. This mechanical trauma not only directly disrupts the anatomical integrity of the aqueous drainage pathways but also triggers chronic inflammatory responses and fibroblastic repair mechanisms, subsequently altering the microenvironment of the angle tissues[
1]. Clinical observations indicate that approximately 20% to 50% of patients with angle recession will gradually develop secondary open-angle glaucoma over months or even decades following the trauma, a process typically characterized by its insidious and progressive nature[
2]. The sustained elevation of intraocular pressure inflicts irreversible damage on optic nerve fibers, eventually resulting in visual field defects or complete blindness, positioning this complication as the second leading cause of irreversible blindness globally.
Imaging techniques play a pivotal role in the clinical diagnosis of angle recession[
3]. Although optical coherence tomography (OCT) provides high-resolution cross-sectional images of the retina, its penetration into the deep structures of the anterior segment is limited; this makes it difficult to accurately evaluate angle morphology, particularly in the presence of common post-traumatic complications such as corneal edema or hyphema[
4]. Ultrasound biomicroscopy (UBM), utilizing high-frequency ultrasound probes (10–50 MHz), overcomes the limitations of optical imaging to achieve 360° panoramic scanning. It clearly visualizes critical pathological features, such as the extent of iridodialysis and the degree of ciliary body detachment, with an axial resolution of 20–60 μm. This provides clinicians with diagnostic evidence nearly equivalent to histology[
5], establishing UBM as the gold standard for diagnosing this condition.
Recent studies have demonstrated the value of artificial intelligence (AI) for medical image analysis. Liu et al.[
6] and Yang et al.[
7] reviewed AI-assisted diagnosis for glaucoma based on multimodal data. Yang et al.[
8] addressed the challenge of identifying subtle lesions in macular degeneration by introducing self-attention mechanisms. Furthermore, Fan et al.[
9] validated the efficacy of deep learning for the detection and grading of diabetic retinopathy.
The aforementioned analysis demonstrates that while intelligent imaging technology has achieved significant progress in the identification of ocular diseases such as glaucoma, macular degeneration, and retinopathy, automated identification research specifically targeting traumatic angle recession remains underexplored. Within the existing body of research regarding angle identification, current outcomes are largely confined to the classification of open/closed angle states and the measurement of specific angle parameters[
10–
12]. Accordingly, the present study focuses on two core objectives: first, the construction of a high-quality human angle recession UBM image dataset to provide reliable data support for model training; and second, the development of a specialized identification model for angle recession based on UBM images to provide a feasible technical pathway for the clinically assisted screening and intelligent diagnosis of traumatic angle recession.
Materials and methods
YOLOv8 model principles
YOLOv8 represents a significant iteration within the YOLO family of object detection algorithms, substantially enhancing detection accuracy while maintaining high inference speeds[
13]. As illustrated in Figure 1, the network architecture of YOLOv8 primarily consists of three key components: the Backbone, the Neck, and the Head. The Backbone is responsible for extracting rich feature information from input images, employing a series of convolutional and deconvolutional layers integrated with residual connections and bottleneck structures. These designs not only facilitate the capture of multi-level image features but also effectively mitigate the gradient vanishing problem in deep networks, ensuring the model can efficiently handle object detection tasks in complex scenarios. The Neck incorporates a Path Aggregation Network-Feature Pyramid Network (PAN-FPN) structure, which enhances detection capability for targets at various scales by fusing feature maps from different stages of the Backbone. The implementation of PAN-FPN is a prominent feature of YOLOv8, distinguishing it significantly from previous versions of the YOLO algorithm series. This multi-scale feature fusion mechanism enables the model to achieve more precise detection across objects of various sizes, making it particularly suitable for scenarios containing with small or densely packed targets. The Head is responsible for the final object detection and classification tasks, comprising localization and classification components. YOLOv8 introduces a decoupled head design, which splits the previously unified detection head into independent branches for detection and classification. This design optimizes the task-specific feature extraction paths for each objective, thereby improving the overall performance of the model. Specifically, the detection head focuses on bounding box localization and regression, while the classification head concentrates on target category recognition, resulting in enhanced detection precision and robustness.
YOLOv8 sub-models
The present study employed five YOLOv8 sub-models—YOLOv8‑nano (YOLOv8n), YOLOv8‑small (YOLOv8s), YOLOv8‑medium (YOLOv8m), YOLOv8‑large (YOLOv8l), and YOLOv8‑extra‑large (YOLOv8x)—to achieve intelligent identification of angle recession, with these variants differing in parameter count, model scale, and computational complexity. YOLOv8n is the most lightweight variant, containing only 3.16 M parameters; its streamlined architecture enables rapid execution on low-computational-resource devices, processing hundreds of images per second, making it ideal for scenarios requiring extreme detection speeds under limited hardware resources. YOLOv8s, with 11.17 M parameters, features an expanded network structure compared with YOLOv8n by increasing the number of convolutional layers and channels, thereby learning richer features and significantly enhancing detection precision while maintaining high efficiency in standard desktop environments. YOLOv8m reaches a parameter count of 25.90 M and utilizes a more complex architecture with strengthened feature extraction capabilities, making it suitable for applications requiring higher accuracy where computational resources are relatively sufficient. YOLOv8l, with 43.69 M parameters, further deepens and widens the network to process more intricate image features, delivering superior performance on large-scale datasets at the cost of increased computational overhead and inference latency. Finally, YOLOv8x is the largest model with 68.23 M parameters, possessing the most robust feature extraction capabilities and a sophisticated network architecture; it is specifically designed for scientific research and high-end clinical diagnostic scenarios where detection precision is paramount and substantial computing power is available.
Construction of the ocular UBM image dataset
Data acquisition
Due to the current scarcity of publicly available human UBM image datasets for angle recession, the training and validation of relevant automated identification algorithms have been significantly hindered. To address this, this study prioritized the construction of a high-quality, specialized dataset to establish a robust foundation for subsequent algorithmic development. The ocular UBM images utilized in this research were obtained from clinical practice at Peking University Third Hospital, and the study was conducted in strict accordance with the Declaration of Helsinki. During data acquisition, the research team systematically collected ocular UBM images using professional UBM equipment from patients across diverse age groups, genders, and clinical backgrounds to ensure sample diversity and clinical representativeness. Following rigorous screening and integration, a total of 212 high-quality ocular UBM images were acquired, comprising 90 images of normal angles and 122 images of angle recession.
Data augmentation
To further increase the volume and diversity of the dataset while simulating clinical conditions, where UBM image acquisition is often compromised by inherent ultrasonic speckle noise and variations in saline or gel coupling that result in reduced signal-to-noise ratios or localized illumination inconsistencies, four data augmentation techniques were applied to the clinically collected images: random brightness enhancement, random brightness reduction, Gaussian noise addition, and salt-and-pepper noise addition. For random brightness adjustment, pixel intensity values were stochastically modified within a range of ± 70 relative to their baseline. To maintain validity within the 8-bit image format, adjusted values exceeding the upper limit of 255 were clipped to 255, while those falling below the lower limit of 0 were set to 0. For the addition of Gaussian noise, a mean of 0 and a standard deviation of 0.3 were utilized to introduce background noise, thereby enhancing model robustness in complex environments. Simultaneously, salt-and-pepper noise was introduced with a density of 0.3 to simulate potential random impulse interference. By applying these four techniques to the original 212 clinical images, a total of 848 augmented UBM images were generated, comprising 360 normal angle images and 488 angle recession images.
Data annotation
Upon completion of the data augmentation, members of the research medical team conducted a comprehensive annotation of both the original clinical and the augmented UBM images. This process was performed using the LabelImg software, employing horizontal bounding boxes to precisely delineate the anterior chamber angle structures and categorize the images as either “normal angle (labeled normal)” or “angle recession (labeled back)”. All annotations were stored in the YOLO format to ensure seamless integration with the subsequent deep learning architectures. The finalized dataset comprised 1060 UBM images, consisting of 450 normal angle images (90 original and 360 augmented) and 610 angle recession images (122 original and 488 augmented). This dataset was partitioned into training and validation sets at a ratio of 7:3. To prevent overfitting arising from the inherent correlation between original and augmented data, a rigorous partitioning protocol was enforced: each original image and its entire suite of augmented derivatives were assigned exclusively to the same set (either training or validation). This strategy reduced the risk of information leakage by ensuring that the model was not evaluated on an augmented version of an image assigned to the training set; however, it did not replace independent external validation. Typical image features post-annotation are exemplified in Figure 2, which displays representative UBM samples for the “normal” and “back” classes.
All images in this study were collected at a single medical center using a single UBM imaging system, and the reported results were obtained from internal validation only. No external-center, multi-device, or temporally independent validation cohort was included.
Model training
The hardware and software environments for model training were configured as follows: Python version 3.9.21, PyTorch version 1.8.0, and Conda version 4.12.0, with computations performed on an AMD Ryzen 7 9700X CPU and an NVIDIA RTX 4070 Ti SUPER GPU. The training hyperparameters were specified with an input image size of 640 × 640 pixels, a total of 200 training epochs, a batch size of 4, and an initial learning rate of 0.01. Following the cosine annealing strategy for learning rate decay inherent to YOLOv8, the final learning rate was reduced to 0.0001 to ensure the overall stability of the network. Figure 3 illustrates the learning curves for the YOLOv8s sub-model, demonstrating that the model reached convergence after 200 training epochs.
Performance evaluation metrics
To assess the model’s performance, metrics commonly used in the field of object detection were adopted, including precision, recall, mAP0.5, mAP0.5:0.95, and giga floating-point operations (GFLOPs). Higher values for precision, recall, and the mAP metrics indicate superior comprehensive performance regarding target recognition accuracy, detection capability, and localization precision, while a lower GFLOPs value represents lower computational complexity and higher inference efficiency. Specifically, precision and recall are utilized to measure the model’s accuracy and detection capability, respectively, and are defined as follows.
In these expressions, true positive (TP) denotes the number of targets correctly identified as the positive class; false positive (FP) represents the number of samples incorrectly identified as positive; and false negative (FN) signifies the number of ground-truth targets that the model failed to detect. mAP0.5 and mAP0.5:0.95 are utilized to comprehensively evaluate both the localization and classification performance of detection boxes under various intersection over union (IoU) thresholds, where IoU represents the degree of overlap between the predicted and ground-truth boxes. In this study, mAP0.5 indicates the mean average precision (mAP) calculated at an IoU threshold of 0.5, whereas mAP0.5:0.95 represents the average of the AP values calculated at multiple IoU thresholds within the range of 0.5 to 0.95. The corresponding calculation formulas are as follows.
Where P denotes the number of selected IoU thresholds, and AP is derived from the integration of the precision-recall curve. GFLOPs is utilized to measure the volume of floating-point operations required during a single forward inference, serving as a critical metric that reflects the computational complexity and inference cost of the model.
Results
Performance comparison of YOLOv8 sub-models
Table 1 summarizes the overall training performance of the five YOLOv8 sub-models (YOLOv8n, YOLOv8s, YOLOv8m, YOLOv8l, and YOLOv8x) on the training and validation sets, while Table 2 details their efficacy in identifying normal angles and angle recession. Based on the data in Table 2, YOLOv8n achieved an mAP0.5 of 0.968 for normal angle identification, but its mAP0.5:0.95 for angle recession was only 0.432, indicating insufficient localization precision and relatively weak overall performance. YOLOv8s exhibited outstanding performance in normal angle identification with a recall of 0.990 and precision of 0.927; for angle recession, its precision and recall reached 0.942 and 0.793, respectively, representing a well-balanced profile among the lightweight models. YOLOv8m demonstrated high precision (0.964) but low recall (0.796) in normal angle identification, while achieving a recall of 0.926 and precision of 0.813 for angle recession, reflecting a distinct performance trade-off between the two classification tasks. The overall performance of YOLOv8l remained relatively stable, with an mAP0.5 of 0.973 for normal angle identification and precision and recall values of 0.910 and 0.911 for angle recession, though it did not exhibit a significant advantage in key indicators.
Overall, the YOLOv8x model demonstrated the most balanced performance among the five sub-models. Although its performance in normal angle identification was surpassed by other sub-models, it achieved the highest precision and mAP0.5:0.95 in angle recession identification, enabling more accurate detection of pathology. The superior performance of YOLOv8x compared to the other four sub-models is primarily attributed to its larger model capacity and more sophisticated architecture. Figure 4a illustrates the identification results of the YOLOv8x model on selected UBM images.
Table 3 presents the identification performance of classical object detection models, such as the Single Shot Multibox Detector (SSD)[
14] and Faster Region-Based Convolutional Neural Network (Faster R-CNN)[
15], alongside the advanced Real-Time Detection Transformer (RT-DETR)[
16]. The results showed that the two traditional models significantly underperformed compared with the five YOLOv8 sub-models. Furthermore, the RT-DETR model achieved an mAP
0.5 of 0.914 in the angle recession detection task; this result further corroborates the targeted effectiveness and superior performance of the proposed improved algorithm in processing the complex features of ocular ultrasound images.
YOLOv8x-CCFM and YOLOv8x-CCFM-SENetV2 models
Although the YOLOv8x model exhibited the most balanced performance among the five sub-models, its capability to identify normal angles remained to be further optimized. To enhance the diagnostic performance of the model, the Cross-Scale Feature Fusion Module (CCFM) and Squeeze-and-Excitation Networks V2 (SENetV2)[
17] were integrated into the YOLOv8x baseline to construct the YOLOv8x-CCFM and YOLOv8x-CCFM-SENetV2 models, respectively; the former incorporates only the CCFM, while the latter integrates both modules.
CCFM is a lightweight feature fusion module designed for object detection tasks, capable of integrating features across different scales through fusion operations to enhance the model’s adaptability to scale variations and its multi-scale object detection capabilities while maintaining a lightweight architecture. The SENetV2 module incorporates multi-branch fully connected layers into the residual block, enhancing the global representation between channels via an aggregated fully connected layer with a cardinality of 4, and subsequently concatenating the multi-branch outputs to restore the feature shape during the excitation stage. Within the YOLOv8 Backbone, the SENetV2 module replaces the original C2f module, aiming to adaptively adjust the channel weights of the feature maps by modeling inter-channel dependencies, thereby emphasizing critical features while suppressing irrelevant information. The architecture of the improved YOLOv8x-CCFM-SENetV2 is illustrated in Figure 5. To address the challenges of varying lesion scales and background interference in UBM images, this study introduces the CCFM to optimize the feature fusion path for enhanced multi-scale expression and integrates SENetV2 to suppress background noise via a channel attention mechanism. The synergy between these two modules significantly improves the model’s ability to capture complex anatomical features and its anti-interference performance, providing a robust structural foundation for high-precision identification.
Experimental results and analysis
The overall training performance of the YOLOv8x-CCFM and YOLOv8x-CCFM-SENetV2 models on the training and validation sets is summarized in Table 1, and their efficacy in identifying normal angles and angle recession is presented in Table 2. According to the data in Table 2, YOLOv8x-CCFM demonstrated limited improvement over the baseline YOLOv8x. However, this improvement was specific to the relatively permissive IoU threshold of 0.5. As shown in Table 2, for the angle-recession class, precision decreased from 0.975 to 0.924, recall decreased from 0.869 to 0.810, and mAP0.5:0.95 decreased from 0.468 to 0.408 compared with the YOLOv8x baseline. Thus, the improved model increased mAP0.5 by 2.1 percentage points but did not consistently improve classification or bounding-box localization at stricter IoU thresholds. This trade-off suggests that the added feature-fusion and attention modules improved coarse detection under IoU = 0.5, while tighter localization of morphologically variable recessed lesions remained challenging. Furthermore, for normal angle identification, the precision, recall, and mAP0.5 improved by 5.60%, 4.00%, and 5.40%, respectively, effectively addressing the suboptimal performance of the original YOLOv8x in the normal angle identification task relative to other sub-models. Regarding the phenomenon where the improved model’s precision for normal angles (0.949) surpassed that for angle recession (0.924), this primarily stems from the highly uniform anatomical structure of normal angles compared to the complex and diverse morphological manifestations of recessed lesions. Following the introduction of the CCFM and SENetV2 modules to bolster feature learning capabilities, the model was able to more precisely localize standardized normal anatomical structures, resulting in a more significant increase in identification precision for these cases than for the morphologically variable pathological samples. Figure 4b illustrates the identification performance of the YOLOv8x-CCFM-SENetV2 model on selected UBM images.
Discussion
The present study evaluated five YOLOv8 model scales and further incorporated CCFM and SENetV2 into YOLOv8x for UBM-based angle recession identification. The improved model achieved better performance for normal-angle identification and increased the angle-recession mAP at an IoU threshold of 0.5 on the internal validation set.
Previous deep-learning studies of the anterior chamber angle have mainly addressed open- versus closed-angle classification or the assessment of angle parameters. The present work focuses specifically on traumatic angle recession and addresses a different clinical identification task. The results should be interpreted as preliminary algorithmic evidence rather than as a replacement for ophthalmologist assessment.
The study is limited by its single-center, single-device design and the absence of external-center, multi-device, or temporally independent validation. Differences in patient populations, imaging protocols, device characteristics, operator experience, and image quality may affect performance. Prospective multicenter validation is required before clinical deployment.
Conclusion
This study established a high-quality human UBM image dataset for angle recession and developed an intelligent identification model based on the YOLOv8 framework. The YOLOv8x-CCFM-SENetV2 model, derived from the YOLOv8x architecture, achieved identification precisions of 0.924 for angle recession and 0.949 for normal angles, demonstrating promising performance for angle recession identification on the internal validation set. Future research will focus on expanding the dataset, enhancing cross-device generalization, and investigating efficient deployment and interpretability techniques to facilitate clinical translation in real-world scenarios.
The generalizability of the proposed method remains to be established. Differences in patient populations, disease spectrum, imaging protocols, UBM device characteristics, operator experience, and image quality may affect model performance. Future work should include multicenter, multi-device, and temporally independent validation using prospectively collected patient-level data before clinical deployment.
The Author(s) 2026. This article is published by Higher Education Press at journal.hep.com.cn.