Toward a standardized methodological framework for developing computer vision models in staging laparoscopy
Francesca Tozzi , Seyed Amir Mousavi , Robbe De Muynck , Dario Quintini , Adris Molnar , Femke Van Vaerenbergh , Xander De Lille , Matthias Van Liefferinge , Wim Ceelen , Wouter Willaert , Wesley De Neve , Niki Rashidian
Artificial Intelligence Surgery ›› 2026, Vol. 6 ›› Issue (2) : 300 -19.
Aim: To evaluate deep learning models for anatomical structure and peritoneal metastasis (PM) detection and segmentation during staging laparoscopy (SL) using a phase-independent dataset, and to quantify how annotation strategy and spatial representation relate to predictive performance.
Methods: A checklist covering 25 anatomical structures, one surgical instrument, and PM was defined. Detection models (YOLOv9, Co-DETR) and segmentation models (SegFormer, Mask2Former) were trained under two label configurations. Videos were split at the video level (60/20/20). To quantify annotation distribution and spatial representation, two class-level descriptors were derived from the training set: object count and area fraction (percentage of image area occupied by each class). Class-level associations between these descriptors and test-set performance [F1-score, Intersection over Union (IoU)] were evaluated using Spearman correlation.
Results: Thirty SL videos yielded 2,309 annotated frames (1,304/433/572 for training/validation/testing). YOLOv9 reached mean mAP@50 of 0.52 and 0.61; Mask2Former achieved mean IoU of 0.51 and 0.61 and F1-scores of 0.65 and 0.73 for Sets A and B, respectively. Despite 4,094 annotations, PM remained difficult to segment (IoU 0.29-0.30; F1-score 0.45-0.46), due to low area fraction and high heterogeneity. For IoU, area fraction showed stronger correlations with performance than object count (ρ up to 0.66 vs. 0.48). Similar differences were observed for F1-score.
Conclusions: Anatomical detection and segmentation during SL are feasible but limited by small-target representation and heterogeneous intra-abdominal context. Spatial representation is more closely associated with segmentation performance than annotation frequency, supporting annotation strategies that address sparse pixel coverage in phase-independent intra-abdominal models.
Anatomical detection / anatomical segmentation / annotation strategy / computer vision / deep learning / laparoscopic staging / peritoneal metastases
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
Salort-Benejam L, Agudo A. NeRFscopy: neural radiance fields for in-vivo time-varying tissues from endoscopy. arXiv 2026;arXiv:2602.15775. Available from https://arxiv.org/abs/2401.00044 [accessed 27 May 2026]. |
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
Sener O, Savarese S. Active learning for convolutional neural networks: a core-set approach. arXiv 2017;arXiv:1708.00489. Available from https://doi.org/10.48550/arXiv.1708.00489 [accessed 27 May 2026]. |
| [29] |
|
| [30] |
Zong Z, Song G, Liu Y. DETRs with collaborative hybrid assignments training. arXiv 2022;arXiv:2211.12860. Available from https://doi.org/10.48550/arXiv.2211.12860 [accessed 27 May 2026]. |
| [31] |
Xie E, Wang W, Yu Z, Anandkumar A, Alvarez JM, Luo P. SegFormer: simple and efficient design for semantic segmentation with transformers. arXiv 2021;arXiv:2105.15203. Available from https://doi.org/10.48550/arXiv.2105.15203 [accessed 27 May 2026]. |
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
Yengera G, Mutter D, Marescaux J, Padoy N. Less is more: surgical phase recognition with less annotations through self-supervised pre-training of CNN-LSTM networks. arXiv 2018;arXiv:180508569. Available from https://doi.org/10.48550/arXiv.1805.08569 [accessed 27 May 2026]. |
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
/
| 〈 |
|
〉 |