1 INTRODUCTION
Surgical training is evolving from Halsted's traditional “see one, do one, teach one” approach to a modern, competency-based model[
1–
3]. During this transition, multiple assessment methods have been developed and refined to better evaluate competency. These include manual assessment tools such as GEARS[
4], OSATS[
5], RACE[
6], EASE[
7], DART[
8], as well as newer semi- or fully-automated assessment methods like automated performance metrics[
9–
12]. The advancement of AI has transformed many areas of medicine, including surgical assessment[
13]. With its robust computational power, AI has the potential to analyze surgical procedures at a highly granular level, potentially action by action, and may eventually replace experts in reviewing surgical videos and evaluating technical skills.
Following this trend, surgical gesture classification has been developed and validated as an innovative tool for surgical assessment. A surgical gesture is defined as the smallest independent unit of surgical instrument-tissue interaction, such as a single cold cut or the spreading of scissors[
14]. Surgery comprises thousands of these individual actions. In the context of surgical training and assessment, procedures can be broken down into various steps, each of which can be further subdivided into distinct surgical gestures. For instance, a robotic-assisted radical prostatectomy can be divided into fourteen steps, such as bladder dropdown, nerve-sparing, and vesicourethral anastomosis[
15]. Nerve-sparing alone includes hundreds of surgical gestures, such as fifty cold cuts, five energy cuts, and the placement of a few hem-lock clips[
16].
Evaluating surgery at the gesture level has several important implications. First, it reveals differences between experts surgeons and trainees in their gesture selection, which can be used to guide surgical training[
14,
17,
18]. Second, it allows for comparisons of dissection styles among expert surgeons, providing an opportunity to correlate these styles with patient outcomes[
16]. Lastly, surgical gestures serve as foundational elements for AI models, enabling these systems to understand and learn the nuances of surgical procedures.
Surgical tasks can broadly be classified into suturing and dissection. In this review, we will discuss these two tasks separately, examine existing gesture classification systems, explore their clinical implications, and provide a state-of-the-art overview of automated gesture recognition models. Finally, we will address current challenges in the field and propose potential solutions.
2 SURGICAL GESTURES FOR SUTURING
Suturing involves a sequence of steps, including needle handling, positioning, entry, driving, withdrawal, followed by knot tying (Figure 1)[
20]. Depending on the direction of needle driving, it can be further classified as (a) forehand versus backhand and (b) overhand versus underhand[
21]. Movement with the palm of the hand facing in the direction of the wrist rotation was defined as forehand. Movement with the dorsum of the hand facing in the direction of the wrist rotation was defined as backhand. A gesture with the needle driver holding the needle in the same plane was defined as flush-hand. Depending on whether needle curvature faced up or down each gesture was further divided into overhand or underhand, respectively. Consistent coupling of two gestures was recognized as a combination gesture, such as right forehand over to under or right backhand under to over. If the needle was driven using several gestures without a discernible pattern, it was categorized as a random gesture[
21].
Defining surgical gestures for suturing holds significant clinical value. First, several suturing skill assessment tools have been developed based on suturing gestures. For example, End-to-end Assessment for Suturing Expertise (EASE) evaluated needle handling, entry, driving, withdrawal, suture placement and knot tying at the level of individual suturing gestures of each stitch[
7]. As such, recognizing suturing gestures serves as the foundation for skill assessment and feedback in surgical practice. Second, a tutorial for vesicourethral anastomosis (VUA) was developed by Chen et al. based on needle-driving gesture classification[
21]. The authors observed how expert surgeons drove their needles at different positions and interviewed them for task breakdowns, illustrating that suturing gesture classification can be effectively applied to surgical training.
Automatic gesture recognition in surgical tasks is actively advancing. Currently, various models are benchmarked against the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS), a data set developed by faculty surgeons at Johns Hopkins University[
22,
23]. The defined surgical gestures in JIGSAWS include:
(G1) Reaching for the needle with right hand.
(G2 Positioning the tip of the needle.
(G3 Pushing needle through the tissue.
(G4 Transferring needle from left to right.
(G5 Moving to center of workspace with needle in grip.
(G6 Pulling suture with left hand.
(G7 Pulling suture with right hand.
(G8 Orienting needle.
(G9 Using right hand to help tighten suture.
(G10 Loosening more suture.
(G11 Dropping suture and moving to end points.
(G12 Reaching for needle with left hand.
(G13 Making C loop around right hand.
(G14 Reaching for suture with right hand.
(G15 Pulling suture with both hands.
This suturing gesture classification covers suturing (G1 to G11 excluding G7), needle-passing (G1 to G6, G8, and G11), and knot-tying (G1 and G11 to G15). State-of-the-art approaches use architectures that incorporate temporal information, utilizing models based on image and/or video data alone or integrating multimodal data, such as robotic kinematics or other sensor information[
13,
24].
3 SURGICAL GESTURES FOR DISSECTION AND EXPOSURE
Despite dissection being a significant component of many common surgical procedures, a well-defined semantic vocabulary system for describing individual dissection gestures is still lacking. To address this gap, Ma et al. manually reviewed surgical videos and classified dissection gestures into three categories: single blunt, single sharp, and combination gestures (multiple actions in sequence). In this system, single blunt dissection gestures are classified as spread, peel/push, and hook, while single sharp gestures include cold cut, hot cut, and burn dissect. Combination gestures involve more complex sequences, such as pedicalize (repeated peels), two-hand spread (both hands pushing in opposing directions), and coagulate-then-cut (coagulating both ends of a tissue/vessel before cutting it in the middle)[
14] (Figure 2).
Although surgical gestures are ideally defined at the single-action level, recognizing combination gestures as an independent category can be valuable, particularly when multiple gestures are repeated frequently within a short period (a few seconds) and serve a specific surgical purpose. In addition to dissection gestures, there are supportive gestures such as retraction, coagulation, clipping, and assistant motions, all of which can be performed by either the dominant or non-dominant hand.
Similarly, Chen et al. and Shafiei et al. proposed two distinct dissection gesture classification systems (Table 1). Chen et al. further classified assistant motions into tamponade, aspiration, and wipe[
17], while Shafiei et al. divided coagulation gestures into bipolar and monopolar cautery[
25].
Variations in dissection gesture usage exist across different procedures and tasks. For example, the composition of dissection gestures differs between partial nephrectomy and radical prostatectomy. Additionally, the execution of specific gestures, such as the length of a hook movement, may vary depending on the procedure.
An association between expertise and gesture usage has been reported. Previous studies have shown that experts take less time to complete the same dissection gestures compared to novices[
14,
17], and that super-experts execute the same gestures differently from experts, potentially impacting patient outcomes[
16]. Additionally, the dissection/exposure ratio varies between experts and novices[
17]. These metrics can serve to assess surgical performance and provide direct feedback to surgical trainees.
Automatic recognition of dissection gestures is actively being developed. Kiyasseh et al. established a state-of-the-art computer vision model, which can recog-nize surgical gestures with AUC ranging from 0.62 to 0.97 across different surgeons, hospitals, and procedures[
26]. Shafiei et al. utilized electroencephalogram (EEG) data to classify surgical gestures, achieving an accuracy of 90%-93% in a single-center study[
25].
4 CONTEXTUALIZING SURGICAL GESTURE
Surgical gestures are granular data, which needs to be put into context to carry greater clinical significance. In this regard, several emerging directions are worth exploring.
First is contextualizing surgical gestures into different anatomic locations. For example, in partial nephrectomy renal hilum dissection, there are three different anatomic locations: renal vein, between renal vein and artery, and renal artery. Gesture usage differs among these locations. More peels/pushes are used around the renal vein rather than hot cuts due to the thinner and more delicate walls of the veins[
14]. In nerve-sparing during radical prostatectomy, gestures may occur in the lateral fascia, prostatic pedicle, or posterior plane[
27]. Depending on the anatomy, the same gesture can carry different clinical implications. For example, a hot cut around the pedicle may differ in significance from a hot cut in the lateral fascia. These distinctions present opportunities for further investigation.
Additionally, there is a new concept called action triplet: <
instrument, verb, target>[
28]. Proposed by Nicolas Padoy's group CAMMA, an open data set
CholecT50 data set was released, which includes 50 videos of cholecystectomy annotated with 161 000 instances from 100 triplet classes. Examples of such action triplets include: <
grasper, retract, gallbladder>, <
hook, dissect, cystic-plate>, <
clipper, clip, cystic-artery>, <
scissors, cut, cystic-duct>, <
irrigator, aspirate, fluid>, etc.[
29]. Challenges using this data set were held at MICCAI in 2021 and 2022[
29–
32]. These triplets recognition can serve as fundamental units for building context-aware systems, with the potential to assist surgeons in decision-making and procedural planning.
Gestures can be evaluated as effective versus ineffective versus erroneous. Inouye et al. utilized dry lab videos of robotic dissection to assess the efficacy of dissection gestures[
33]. Ineffective gestures were defined as those that did not produce a meaningful effect on the tissue (e.g., a cold cut that barely touched the tissue). Effective gestures achieved the intended effect in the target tissue, while erroneous gestures disrupted the tissue in a way that did not align with the task objectives[
33]. Surgical errors vary across procedures and represent another dimension for assessing surgical performance. The relationship between surgical gestures and surgical errors is an area that merits further exploration.
Finally, gestures can be bundled together to form surgical tasks. Surgical workflow recognition has been explored through the automatic decomposition of procedural video recordings, breaking down the sequence of events into actions at varying levels of granularity—ranging from gestures to activities and phases (from fine to coarse decomposition)[
13,
20,
34–
36]. For example, in nerve-sparing step, surgical gestures can form different surgical tasks such as releasing neurovascular bundle, isolating structures, or extending plane[
27]. These groupings provide more context, aiding AI models in understanding the broader scope of surgical procedures.
5 APPLICATION OF SURGICAL GESTURES IN SURGICAL TRAINING
Structured learning in surgical training provides a systematic approach to mastering complex procedures by dividing them into distinct steps. In our previous research, we delineated robot-assisted radical prostatectomy and robot-assisted partial nephrectomy into discrete phases, allowing for clearer, step-by-step learning. Now, surgical gestures are emerging as a valuable method for refining technical skills on a more granular scale, enhancing anatomical comprehension, and improving patient outcomes through more precise, standardized training protocols.
At its core, a surgical gesture is the interaction between instruments and tissue. Each gesture encompasses a specific movement or technique, and when sequenced, they reveal a structured “surgical vocabulary” analogous to a DNA sequence, encoding essential procedural information. By dissecting surgical gestures, we gain a novel perspective on technique. Skill variations can be observed with high resolution, providing objective data that were previously difficult to measure. For instance, long-standing debates about whether sharp or blunt dissection is more effective in certain scenarios can now be evaluated with quantitative analysis, as can the choice of suturing technique or dissecting gesture depending on the anatomical or procedural context.
The incorporation of surgical gestures can enhance communication in surgical feedback, enabling mentors to provide more targeted guidance. This approach not only improves training efficiency but also allows trainees to assimilate complex techniques in manageable steps. Additionally, post-hoc gesture analysis empowers trainees to reflect on their performances with precision. When combined with video review, it promotes active learning, transforming video analysis from a passive activity into a dynamic learning process where trainees can study and improve upon each gesture in a structured manner. Additionally, gesture-based frameworks offer an opportunity to record each critical step systematically, contributing to more efficient documentation and potentially assisting in surgical auditing.
Finally, the breakdown of surgical procedures into gestures could also serve as a foundation for future automation in surgery. With advancements in machine learning, automatic gesture recognition is a promising avenue that could standardize performance evaluation and support real-time feedback, ultimately creating a streamlined, data-rich environment for surgical training.
6 CHALLENGES IN SURGICAL GESTURE RESEARCH
While surgical gesture research is a growing field that has gained increased attention, several challenges make progress difficult.
First, defining surgical gestures can be vague, and annotation is often subjective. Although a certain level of inter-annotator variability may be informative—indicating areas where consensus is lacking or clinical practices are not yet well established—excessive subjectivity can lead to inaccuracies and complicate automation. Clear and consistent communication between annotators, especially those with different backgrounds or from various research groups, is essential to establishing a common language. A potential solution is the development of automatic gesture recognition models that reduce subjectivity in annotation, though the automation of surgical gesture recognition may be difficult because of image occluded by blood and variances between surgeons and hospitals.
Second, annotation is time-consuming. Gesture annotation may require frame-by-frame manual review of surgical videos, and annotating a 10-minute video can take several hours. The number of gesture categories further impacts the time required, highlighting the need to balance granularity and efficiency to ensure manageable annotation workflows.
Third, training annotators is challenging. Gesture annotation demands familiarity with various surgical instruments and a basic understanding of surgical contexts. Triplet annotation or grouping gestures into surgical tasks require even deeper surgical knowledge, which may not be easily accessible to all researchers. Building a multidisciplinary team that includes individuals with a surgical background can improve the quality and accuracy of surgical gesture research.
Finally, suitable annotation software for collaborative work is lacking. Ideally, each annotation should be double-checked, a process that would benefit from dedicated, user-friendly software. A few available sofewares are presented in Table 2.
7 CONCLUSION
Surgical gestures represent the fundamental units of analysis in surgical studies, allowing for a detailed and systematic understanding of surgical techniques. While multiple classification systems currently exist, they often overlap yet lack uniformity. Developing a unified system could standardize terminology and improve the reproducibility of research focused on surgical gestures. Both the choice of gestures and the skill with which they are executed are critical in evaluating surgical proficiency and predicting patient outcomes. Additionally, gestures provide valuable insight into expected surgical results, serving as a basis for providing constructive feedback to surgeons. At this granular level of workflow analysis, surgical gestures also lay the groundwork for future advancements, including intelligent context awareness and AI-driven automation in surgery.
2025 The Author(s). UroPrecision published by John Wiley & Sons Australia, Ltd on behalf of Higher Education Press.