1. State Key Laboratory of Robotics and Systems, Harbin Institute of Technology Shenzhen, Shenzhen 518055, China
2. Children’s Hospital of Fudan University, Shanghai 201102, China
3. School of Computer Science, University of Bristol, Bristol, BS8 1TH, UK
xuxiu@fudan.edu.cn
honghai.liu@icloud.com
Show less
History+
Received
Accepted
Published Online
2026-05-27
2026-08-17
2026-09-11
PDF
(3149KB)
Abstract
It is evident that standardized behavioral biomarker analysis is crucial for early screening of children with autism spectrum disorder (ASD). However, conventional screening methods rely heavily on subjective scale-based assessments, while existing machine-assisted approaches often suffer from imprecise event triggers that hinder robust biomarker localization and clinical adoption. To address this challenge, we propose a novel computer-aided framework that facilitates behavioral biomarker analysis by segmenting clinical protocols into meaningful actions. Specifically, our framework re-conceptualizes the screening process by leveraging temporal action segmentation to parse continuous clinician–child interactions into discrete protocol steps, thereby establishing precise temporal windows to analyze the child’s responsive behaviors. At its core, a clinical-adapted skeleton-based action segmentation model accurately divides the clinician’s movements, followed by a rule-based post-processing module that extracts and scores the child’s behavioral biomarkers. We evaluate our framework on a real-world clinical dataset comprising 96 video recordings from 48 subjects across two distinct protocols. The experimental results demonstrate strong consistency with expert clinical assessments, achieving agreement rates of 82.1% and 93.6% on the respective protocols. This work provides a reliable, scalable, and non-intrusive decision-support pipeline for behavior-based ASD screening. By overcoming the efficiency and cost bottlenecks of manual video analysis, it bridges the gap between video analytics and routine clinical workflows, demonstrating strong potential for real-world deployment.
Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by persistent challenges in social interaction, communication, and the presence of restricted or repetitive behaviors [1]. A 2023 report by the Centers for Disease Control and Prevention (CDC) [2] indicates that 1 in 36 eight-year-old children in the United States has been identified with ASD, while the prevalence in the Chinese population is approximately 0.7% to 1.0% [3]. Globally, the prevalence of ASD has been on the rise, placing a significant burden on individuals, families, and healthcare systems. Early and accurate screening is widely recognized as critical for initiating timely interventions, which can substantially improve long-term developmental outcomes for children with ASD [4].
Traditional clinical assessment and diagnosis rely heavily on standardized diagnostic instruments[5], such as the Autism Diagnostic Observation Schedule-Second Edition [6] and Autism Diagnostic Interview-Revised [7]. These methods, while highly regarded and widely adopted in clinical practice, are often time-consuming and depend significantly on the subjective judgment of clinicians. This subjectivity can lead to inconsistencies in assessment, requiring a high level of clinical expertise, which is in serious short supply globally. In many developing countries, such as China, the number of professionals qualified to administer these diagnostic assessments is critically low, with only a few hundred certified clinicians available to serve more than 13 million children, creating a major bottleneck for early ASD detection.
To mitigate these challenges, researchers have developed machine-assisted auxiliary diagnostic methods. These approaches utilize structured social paradigms, where a clinician interacts with a child in a standardized manner to elicit specific social responses [8–10]. These specific social responses, which serve as behavioral biomarkers, can be captured and interpreted using artificial intelligence tools, as evidence for risk level scoring, such as turning their head in response to their name being called, social smile [11], etc. These paradigms have shown promise in increasing objectivity and consistency compared to purely observational methods.
However, a fundamental challenge persists in these machine-assisted systems: the accurate temporal localization of behavioral biomarkers. Most current research [12,13] focuses on the classification of a child’s behavior but largely overlooks the critical problem of precisely determining when that behavior should be assessed [14,15]. As illustrated in Fig. 1, existing paradigms for localizing relevant interaction phases heavily rely on a coarse-grained manual trigger coupled with frame-based analysis. For instance, clinicians often need to use manual trims, physical triggers (e.g., pressing a Bluetooth button), or external events to mark the start of a social stimulus. Alternatively, systems attempt to detect specific signals frame-by-frame using auxiliary uni-modal tools, such as speech recognition for name-calling [16] or hand position localization for toy requests [17]. And these traditional motion feature extraction and action recognition approaches neglect the high-level temporal context across action frames. Consequently, as depicted in the left panel of Fig. 1, traversing frames blindly without macro-level temporal understanding inevitably introduces severe motion noise, resulting in unstable analysis windows and discrete, fragmented predictions.
To overcome these limitations, we explore the applicability of temporal action segmentation (TAS) in fine-grained parsing of clinician screening procedures, resulting in a novel screening framework named temporal action segmentation-based autism early screening framework (TAS2F). TAS technique aims to partition a long, untrimmed motion sequence into distinct, contiguous action segments and classify each segment. TAS2F applies TAS to automatically and accurately divide the entire screening session into its constituent protocol steps (e.g., “presenting stimulus”, “observing response”). Then, the post-processing module will detect and score the child’s behavioral biomarker following the protocol’s predefined rules. Our framework offers several key advantages: it provides a rich, temporal understanding of the clinician’s workflow, eliminates the need for auxiliary detection models, yields precise action boundaries for subsequent analysis, and can even be used to verify whether the clinician is adhering to the standardized protocol. To further improve the action segmentation performance, we also customize a skeleton-based TAS algorithm, language-assisted part-level segmentation (LAPS-AS), for ASD early screening scenarios.
Our contributions are threefold:
(1) We propose a novel analytical framework TAS2F for ASD early screening based on TAS, shifting the focus from simple event detection to a holistic, temporal understanding of the clinical interaction.
(2) We construct a new, richly annotated dataset for clinician behavior segmentation, comprising 96 recordings from 48 subjects across two distinct ASD screening protocols, to facilitate future research.
(3) We propose a complete pipeline for assisted biomarker analysis, which integrates a clinical-tailored skeleton-based action segmentation model, LAPS-AS, with a rule-based post-processing module.
The remainder of this paper is structured as follows. Section 2 reviews prior work on machine-assisted autism screening and TAS. Section 3 delineates the proposed framework, encompassing the unified clinical protocol workflows [for expressing needs with pointing (ENP) and responding to joint attention (RJA)], the skeleton-based action segmentation algorithm and the downstream scoring criteria. Section 4 outlines the dataset construction and implementation, followed by a comprehensive evaluation of the LAPS-AS segmentation performance and framework consistency. Section 5 discusses our findings, architectural efficacy, bottlenecks, and clinical implications, while Section 6 concludes the paper.
2 Related Work
Our work is positioned at the intersection of machine-assisted autism screening and TAS. In this section, we review key developments in both fields to contextualize our contributions.
2.1 Machine-assisted autism screening
The pursuit of objective biomarkers for ASD has spurred the development of technology-driven screening tools. Some research has focused on neuroimaging and physiological signals like Magnetic Resonance Imaging (MRI) [18] and Electroencephalography (EEG) [19]. Driven by recent advances in computer vision, video-based behavioral quantification has emerged as a prominent method for automated screening, offering a non-invasive, cost-effective, and highly accurate alternative to traditional modalities. To precisely capture diagnostic biomarkers in controlled environments, a significant body of literature has integrated these vision-based techniques with structured behavioral protocols. For example, the respond-to-name (RTN) [20] protocol analyzes a child’s head-turning response to their name being called. Respond-to-introduction (RTI) [17] protocol is designed to evaluate a child’s responsiveness to social cues, characterized by the behavior of delivering a toy to the therapist upon receiving verbal and nonverbal instructions. Other protocols [21,22] focus on joint attention—a core deficit in ASD—by tracking a child’s gaze as they follow a clinician’s point or gesture. Similarly, the ENP [23,24] protocol assesses the quality of a child’s pointing gesture when requesting an object. These protocols typically follow a standardized pipeline: a clinician provides a social stimulus, and the child’s reaction is recorded and analyzed to extract quantitative biomarkers. This approach has demonstrated high accuracy and consistency with traditional scale-based assessments.
Despite their success in standardizing stimulus and response analysis, these methods critically depend on the accurate temporal localization of protocol steps. The current reliance on manual video trimming or simplistic, often unreliable, event triggers represents a significant bottleneck, limiting the scalability and reliability of these promising screening pipelines. This highlights a pressing need for a more robust method for segmenting the interaction flow.
2.2 Temporal action segmentation
TAS [25–28] aims to densely classify every frame in a long, untrimmed video, partitioning it into meaningful, contiguous action segments. This technology is crucial for applications requiring detailed workflow analysis, such as surgical robotics, industrial assembly, and human-computer interaction. Early TAS methods [29] often relied on Red-Green-Blue (RGB) video and handcrafted features, but recent advancements have been dominated by deep learning approaches.
Within TAS, skeleton-based methods (STAS) [30,31] have gained traction due to their robustness to background noise, lighting variations, and inherent privacy-preserving qualities. These methods typically use graph convolutional networks (GCNs) [32–34] to model the spatial relationships between body joints and temporal models like temporal convolutional networks (TCNs) [35–37] or Transformers [38] to capture motion dynamics over time [39–46]. However, many STAS models were originally designed for general human activity recognition and often struggle with the fine-grained and subtle actions characteristic of clinical interactions. They tend to generate “over-smoothed” features [47] that average out the small, discriminative movements—like a subtle hand gesture or a quick glance—that are diagnostically significant in an ASD context.
While action segmentation has proven effective for analyzing structured procedures in various domains [48–51], its application to the nuanced, interactive dynamics of clinical behavioral protocols remains largely unexplored. There is a clear gap in applying this powerful technology to address the critical localization problem in machine-assisted ASD screening. The existing limitations of STAS in capturing fine-grained motion and understanding semantic context further underscore the need for a specialized approach tailored to the unique demands of this clinical application.
3 Method
3.1 Unified modeling and preliminary protocol
Our framework is designed to analyze a class of machine-assisted screening paradigms that, while targeting different behavioral biomarkers, share a common, unified structure. We conceptualize these protocols as a sequence of three core phases and one rule-based scoring engine, as shown in the upper part of Fig. 2.
(1) Engagement phase: The clinician interacts with the child to establish rapport and create a comfortable environment (e.g., playing with a toy).
(2) Stimulus phase: The clinician performs a standardized action to elicit a specific social response from the child (e.g., pointing to an object, placing a desired toy out of reach).
(3) Observation phase: The clinician ceases their action and waits, creating a clear temporal window to observe and evaluate the child’s reaction to the stimulus.
(4) Scoring engine: Define the different reaction types and corresponding scores.
We instantiate our framework across two distinct machine-assisted screening protocols: ENP and RJA.
① ENP: Pointing, a crucial early communicative gesture in which an infant extends their index finger, serves both to express needs (protoimperative pointing) and to direct another’s attention for social sharing (protodeclarative pointing). The emergence of this skill is a milestone in cognitive, linguistic, and social development, and its absence or delay is a significant risk marker for developmental disorders like ASD. Studies have established a strong correlation between early pointing skills and later language abilities. Echoing this, a pediatric consensus highlights the lack of appropriate gestures—particularly purposeful pointing to request items of interest—as a hallmark of communication deficits in ASD. This deficit can manifest as a reduced frequency of overall gesture use in children as young as 12 months. The ENP protocol workflow is as follows:
1) Engagement phase: The clinician and child are seated opposite each other at a table. The clinician engages the child by playing with a toy, such as blowing bubbles, to capture their interest and encourage interaction (e.g., having the child pop the bubbles).
2) Stimulus phase: After a period of play, the clinician makes the toy inaccessible by closing the bubble bottle and placing it in a location that is visible but out of the child’s reach. This process is repeated with a second toy (e.g., a light-up ball) to provide another trial.
3) Observation phase: The clinician ceases action and observes the child to see if they will use a pointing gesture to request the toy. The core observation is whether the child points to the object to communicate their desire. The clinician may use verbal or physical cues to prompt the child.
4) Scoring engine: Scoring is determined by the type of gesture the child uses to express their need. For the ENP protocol, the analysis occurs during the “observe child’s requesting behavior” windows.
② RJA: Joint attention is a crucial coordinated attentional skill in early socio-cognitive development. It signifies the ability of an infant or child to share a focus on an object or event with another person during a social interaction. This skill, an essential component of early social communication, is typically manifested through behaviors such as eye contact, pointing, or showing objects. Joint attention comprises two key components: initiating joint attention (IJA) and RJA. IJA occurs when a child actively guides another person’s attention to an object or event. Conversely, RJA refers to a child’s ability to react to another’s directional cues, such as following their pointing gesture or gaze. This ability is particularly significant in the early social development of children with autism, as research indicates they exhibit marked deficits in joint attention. The RJA protocol workflow is as follows:
1) Engagement phase: The clinician engages the child through play to establish rapport and gain their attention. The child is seated at a table, often with a parent present for comfort, while the clinician uses toys and games to help the child acclimate to the environment.
2) Stimulus phase: The clinician points to a target object located outside the child’s immediate field of view to elicit a response.
3) Observation phase: The clinician observes whether the child follows the pointing gesture to look at the target object. The core observation is whether the child correctly turns their head to follow the clinician’s line of sight. Verbal prompts may be used to guide the child’s attention.
4) Scoring engine: The scoring is based on the child’s response. To apply the framework to the RJA protocol, the key observation window is the “point to a target” step.
3.2 Framework overview
Our proposed framework is illustrated in Fig. 2.
(1) Multimodal data synchronized recording: We utilize an audio-visual acquisition system to record multimodal data throughout the entire clinical paradigm. Within this pipeline, continuous human skeleton sequences are leveraged for TAS, while appearance-based character and scene object clues are utilized to evaluate behavioral biomarkers. The specific processing modalities are strictly determined by the downstream requirements of action stream matching and biomarker scoring; for instance, the acoustic stream is utilized to capture event-triggering signals, whereas the raw RGB frames are adopted for the child’s head pose estimation.
(2) Identity identification and feature extraction: During the preprocessing and feature extraction stage, we first identify the clinician and the child in the video based on their relative positional relationships; specifically, the child always sits on the left side of the interaction table, while the clinician sits on the right side. Subsequently, taking the raw video as input, a 2D pose estimation algorithm is applied to extract the clinician’s skeleton sequence. This sequence is mathematically formulated as a spatiotemporal tensor:
where , , and denote the spatial coordinates, body joint indices, and continuous temporal frame lengths, respectively. This tensor serves as the direct input to our skeleton-based TAS model LAPS-AS.
(3) Clinician action stream parsing and observation window localization: Taking the clinician’s skeleton sequence as input, the LAPS-AS model executes frame-by-frame action probability estimation:
where represents the predefined set of clinical action classes and denotes the model parameters. Through this classification, the continuous session is automatically partitioned into clinical protocol steps. Specifically, the observation window is instantiated as the “Observe child” action segment for the ENP and the “Point to a target” segment for the RJA, which can be localized as:
(4) Child behavior computation: To analyze the child’s behavior within the localized observation window , our framework utilizes a toolkit of specialized computer vision models.
Person and hand detection: We use RTMDet [52], a high-performance real-time object detector, to locate the child and clinician in the scene. Once the child’s hands are localized, we use a fine-tuned ResNet-50 [53] classifier to recognize the gesture type (e.g., index-finger pointing, whole-hand pointing, neutral).
Gaze and head pose estimation: A hybrid approach is used to determine the child’s direction of attention. For large head movements, we employ the WHENet [54] model, which robustly estimates head pose (yaw, pitch, roll) from a single image. For more subtle gaze shifts where the face is clearly visible, a ResNet-50 based model pre-trained on the ETH-XGaze [55] dataset provides finer-grained eye gaze vectors.
General object detection: To address the challenges in screening paradigms that utilize diverse props like children’s toys, we move beyond training specialized object detectors for each context. Such an approach is often infeasible as it suffers from poor scalability and generalization. Instead, our framework integrates Grounding-DINO [56], a zero-shot object detector. By processing natural language prompts, our system can localize any specified object, such as a “toy truck” without the necessity of retraining or fine-tuning the detection model.
(5) Rule-based scoring: Based on this segmentation, a post-processing module analyzes the child’s behavior only during the relevant “observation” windows to extract clinical biomarkers (e.g., pointing gestures, gaze shifts). Finally, it generates a quantitative score based on the protocol’s rules, providing an objective assessment to support the clinician’s evaluation.
This pipeline removes the need for manual video trimming or unreliable physical triggers, ensuring a consistent and efficient analysis of every screening session.
The technical core of our framework is the LAPS-AS model, as shown in Fig. 3. It is specifically designed to address the unique challenges of clinical behavior analysis through two key components: a disentangled part motion encoder (DPE) for capturing fine-grained movements and a language-assisted distribution alignment (LDA) strategy for embedding clinical knowledge into the model.
3.3.1 Disentangled part motion encoder (DPE)
Clinical logic: In a clinical setting, a clinician’s actions are often defined by the subtle movements of specific body parts—a pointing index finger, a head turn, or the coordinated motion of both hands. Standard models that analyze the “whole body” as one unit can miss these crucial details. The DPE is designed to overcome this by modeling the motion of different body parts (e.g., arms, head, torso) independently and in parallel, ensuring that fine-grained, clinically relevant cues are preserved.
Technical implementation: The DPE first extracts a rich bank of spatial features from the input skeleton sequence using a multi-scale graph convolution (MS-GCN) [57–59]. From this, it generates separate feature banks for predefined body parts and for the whole body. These features are then processed in parallel streams using efficient linear-transformer layers to model their temporal evolution. A key process is the part-global interaction module, which uses cross-attention to allow each part-specific stream to be aware of the overall body context, enabling the model to understand both independent part motions and their coordination. The final motion representation is a comprehensive fusion of these part-level and global features.
3.3.2 Language-assisted distribution alignment (LDA)
Clinical logic: A human observer understands the inherent logic of a clinical protocol—for instance, that “blowing bubbles” and “playing with a ball” are both forms of “engaging the child”. Traditional models, trained with simple numeric labels, are blind to these semantic relationships. The LDA strategy addresses this by using the textual descriptions from the clinical protocol to teach the model these relationships. This helps the model learn a more robust and clinically meaningful representation of the actions.
Technical implementation: We use a large language model (LLM) like GPT-4 to generate rich textual descriptions for each action class in the protocol. These descriptions are encoded into high-dimensional vectors using a pretrained text encoder (e.g., CLIP), creating a “semantic map” where similar actions are located closer to each other. During training, we introduce a novel loss function, the Skeleton-Text Alignment Loss (), based on Kullback-Leibler (KL) divergence. This loss function encourages the model to map the motion features of an action to the corresponding location on the semantic map. This alignment process forces the model to structure its internal representations according to the clinical logic embedded in the language, leading to better intra-class compactness and clearer decision boundaries between different actions.
3.3.3 Semantic offset adapter (SOA) and final model
To account for the natural gap between abstract text descriptions and concrete physical movements, we include an SOA. This is a small neural network that learns to slightly adjust the positions of the text-based “semantic anchors” during training, allowing for a more flexible and accurate alignment with the real-world motion data.
The features from the DPE are passed to a temporal refinement module, which includes an action segmentation branch and a boundary regression branch. The latter is specifically trained to predict the precise start and end frames of each action, further improving segmentation accuracy. The model is trained with a composite loss function that combines standard classification losses with our novel alignment loss: . The language-assisted components (LDA and SOA) are only used during training to guide the learning process and are discarded during inference, adding no computational overhead to the final deployment.
3.3.4 Post-processing and scoring mechanisms
ENP Protocol: The biomarker involves hand gesture, pointing direction, and eye contact with the clinician. The rule engine processes the parsed child behaviors as follows:
• Score 0 (complete pointing): Awarded if the toolkit detects an “index-finger pointing” gesture AND the pointing vector intersects the toy’s bounding box AND the child’s gaze vector intersects the clinician’s bounding box.
• Score 1 (incomplete pointing): Awarded if (the toolkit detects “index-finger pointing” at the toy but without eye contact) OR (it detects “whole-hand pointing” at the toy WITH eye contact).
• Score 2 (no pointing): The default score if neither of the above conditions is met within the observation window.
The detailed scoring mechanism of the ENP protocol is formulated in Algorithm 1.
RJA Protocol: The biomarker criterion is whether the child’s gaze aligns with the target object. The rule engine checks the child’s gaze vector frame-by-frame. Since the target’s location is fixed, this simplifies to checking if the gaze angle falls within a predefined range (Pitch > −35°, Yaw between −30° and 20°).
• Score 0: if the child correctly follows the point to the object, indicating a successful joint attention response.
• Score 1: if the child fails to follow the point.
The detailed scoring mechanism of the RJA protocol is formulated in Algorithm 2.
4 Experiments
4.1 Dataset construction
We recruited 48 children (25 diagnosed with ASD, 23 typically developing) at the Children’s Hospital of Fudan University, with an average age of 27.1 months. Each child participated in both the ENP and RJA protocols, resulting in 96 original video recordings. Of the 96 original recordings, 95 ENP sessions and 94 RJA sessions were included in the consistency analysis, as three sessions were excluded due to the children moving out of the camera view or incomplate task execution. The total dataset comprises 231.72 minutes of interaction video (347,580 frames). All procedures were approved by the hospital’s IRB (No. 2019028). The recordings were conducted on a standardized platform equipped with an Azure Kinect camera capturing at 25 fps. The video frames were then manually annotated by experts according to the action definitions (see in the Appendix).
4.2 Framework performance evaluation settings
To assess the clinical validity of our TAS2F framework, we conducted algorithm performance evaluation and two consistency analyses.
(1) Algorithm performance: We utilize the common metrics utilized in action segmentation task, to report the performance of the proposed LAPS-AS against several state-of-the-art skeleton-based action segmentation methods. We report three commonly used metrics for temporal action segmentation: frame-wise accuracy (Acc), segmental Edit score, and F1-score at different temporal IoU thresholds (F1@k).
(2) Segmentation consistency: We evaluate the consistency between the two segmentation approaches: the algorithm-based and the manually-annotated. The final scores generated by our framework using action segmentation against the scores generated using ground-truth segmentation are compared to investigate the applicability of the TAS module.
(3) Scoring consistency: We compared the scores generated by our framework (using manual segmentation) against the original scores provided by expert clinicians to evaluate the effectiveness of the whole framework.
4.3 Implementation
Our model was implemented in PyTorch. We used a six-fold cross-validation scheme, splitting the 48 subjects into a training set of 38, a validation set of 2, and a test set of 8. We used the Adam optimizer with a learning rate of 0.0005 and trained for up to 200 epochs on an NVIDIA RTX 3090 GPU.
4.4 Algorithm performance
ENP protocol: As shown in Table 1, LAPS-AS achieves the best performance on the ENP dataset, particularly in the F1 and Edit scores. The ENP protocol involves nine action classes with high inter-class similarity (e.g., placing the bubble bottle vs. placing the light ball). The superior performance of LAPS-AS highlights the effectiveness of its DPE in capturing subtle, part-level differences in hand grasp and arm movement, while the LDA helps the model distinguish between semantically similar but distinct protocol steps.
RJA protocol: The RJA protocol presents a distinct operational challenge: it consists of only two classes (“Background” and “Point to target”), where the target pointing action is highly transient, making precise boundary detection critical. As summarized in Table 2, our LAPS-AS framework consistently outperforms all competitive baselines across all evaluated metrics. Our method achieves the highest accuracy of 99.0% and an Edit score of 95.4%. More importantly, LAPS-AS establishes a substantial performance margin in both F1@0.1 and F1@0.5, reaching 96.8%. This exceptional robustness under stricter temporal overlap constraints directly demonstrates the model’s capacity to precisely localize the onset and offset of short-duration actions. Accurate boundary detection at this level is crucial for ensuring that downstream clinical screening and child response analysis are anchored within the exact valid temporal window.
4.5 Segmentation consistency and scoring consistency
The results are summarized in Table 3. For segmentation consistency, our framework achieved a high agreement of 91.6% for ENP and 95.7% for RJA with the scores derived from manual segmentation. This demonstrates that our action segmentation model is robust and accurate enough for practical clinical use, effectively replacing the need for manual video trimming.
For scoring consistency, the framework’s scores showed 93.6% agreement with expert scores on the RJA protocol, indicating that our gaze analysis module is highly effective. The agreement on the ENP protocol was lower at 82.1%. The primary source of error was the child perception module, which sometimes struggled with challenges like hand occlusion, ambiguous gestures, and rapid gaze shifts, making the multi-faceted ENP scoring more difficult than the RJA task.
4.6 Ablation on main modules
The results of the ablation experiment are shown in Table 4.
Effect of DPE: Integrating the DPE module into the baseline improves segment-level metrics. This validates our clinical intuition: by decoupling localized body parts (such as pointing index fingers and head turns) from global skeletal movement via part-global cross-attention, DPE effectively isolates fine-grained diagnostic cues and mitigates over-segmentation.
Effect of LDA: Adding the LDA strategy yields the most substantial performance leap across all metrics, lifting the frame-wise accuracy to 86.2% on ENP and 98.6% on RJA. This confirms that guiding the skeletal feature manifold using LLM-derived clinical knowledge via KL-divergence alignment enforces tighter intra-class compactness and establishes clearer decision boundaries between semantically overlapping action classes.
Effect of SOA: Finally, incorporating the SOA brings the framework to its optimal performance across both protocols (achieving 86.6% Acc/92.0% F1@0.1 on ENP, and 99.0% Acc/96.8% F1@0.1 on RJA). This demonstrates that SOA effectively bridges the remaining domain gap between abstract textual anchors and continuous dynamic skeletal embeddings.
4.7 Qualitative results and visualization
To qualitatively evaluate the performance of our proposed method in TAS, we visualize the prediction completion maps against several competitive baselines [including MS-GCN, Action Segment Refinement Framework (ASRF), and Decoupled Spcdio-Temporal Framework (DeST)] along with the ground-truth annotations across both ENP and RJA screening protocols.
As illustrated in Fig. 4, for the complex ENP protocol which contains dense and highly contiguous behavioral phases, traditional baseline models suffer severely from over-segmentation and false positives. Specifically, MS-GCN and DeST introduce considerable noise during the transition phases (e.g., between the bubble blowing and observation states), occasionally resulting in fragmented action blocks or missing brief actions entirely. In sharp contrast, our framework produces highly coherent segment boundaries that closely mirror the ground truth. This superior alignment demonstrates that incorporating the LDA strategy successfully injects clinical constraints, ensuring that the model differentiates ambiguous action boundaries via global semantic priors.
Furthermore, Fig. 5 presents the visual comparisons on the RJA protocol. Given that the RJA paradigm is characterized by a sparse execution flow dominated by the “Background” state and transient “Point at target” actions, capturing the exact temporal onset of the pointing gesture is highly challenging. Our model precisely anchors the start and end frames of the pointing behavior without generating erratic predictions in the extended background sequences. Collectively, these qualitative visualizations offer compelling evidence that our cross-modal representation scheme significantly strengthens intra-class compactness and clears up decision boundaries, thereby yielding robust behavioral biomarker localization for clinical developmental screening.
4.8 Confusion matrix analysis
As shown in Fig. 6, (A) and (B) present the confusion matrices for the RJA and ENP protocols, respectively. For the RJA protocol, the model demonstrates robust overall classification capability, with minor confusions mainly occurring at action boundary shifts between “Background” and “Point at target”. For the ENP protocol, the model achieves strong overall prediction performance, with primary confusions concentrated between “Background” and “Observe child”. This is mainly because clinicians occasionally exhibit brief irrelevant behaviors during observation, such as turning around to talk with others or recording child behaviors, which can easily be misclassified as background. Other minor confusions are predominantly located near action transition boundaries, such as between “Open bubble bottle” and “Blowing bubbles”, as well as “Play the ball with child” and “Observe child”.
5 Discussion
5.1 Advantages of the proposed method
Demonstrating the feasibility and robustness of a computer-aided pipeline, this study introduces a conceptual paradigm shift in the field—moving away from isolated, event-trigger-based detection of behavioral windows towards a holistic, process-oriented understanding of the entire clinical interaction. This transition is essential as it establishes a more comprehensive and stable analytical framework. By segmenting the clinician’s workflow into semantically meaningful steps, our approach provides a clear temporal scaffold upon which diverse analytical toolkits—including those for child behavior parsing—can be automatically and coherently integrated. Consequently, this transforms the analytical workflow from a series of disjointed detection tasks into a unified, context-aware system.
5.2 Clinical significance
Our framework achieves high segmentation consistency scores of 91.6% and 95.7%, proving that a computer-aided pipeline can serve as a reliable alternative to laborious and time-consuming manual video annotation. In clinical practice, this breakthrough offers dual value: on one hand, it enables high-throughput analysis of large-scale video data, significantly accelerating research progress and breaking the bottleneck of costly and slow manual processing; on the other hand, it paves the way for scalable clinical deployment. By allowing objective, behavior-based screening to be seamlessly integrated into routine clinical workflows without increasing the burden on healthcare professionals, it directly addresses a core bottleneck in the accessibility of early ASD screening.
In the long term, the application of this technology extends far beyond initial screening. First, it can function as an objective training and quality assurance tool for clinicians, offering quantitative feedback on adherence to standardized protocols during training and certification. Second, it can be employed for longitudinal monitoring of intervention efficacy by tracking changes in the frequency and quality of children’s social-communicative behaviors (such as joint attention and requesting gestures) to provide objective, data-driven metrics for developmental progress. Finally, when applied to large-scale datasets, it serves as a powerful engine for digital phenotyping, assisting researchers in discovering novel, subtle, and quantifiable behavioral biomarkers that may be difficult for human observers to detect.
5.3 Evaluation of failure causes
The discrepancy in scoring consistency between the RJA protocol (93.6%) and the ENP protocol (82.1%) objectively reveals that the system’s primary performance bottleneck lies in the complex downstream parsing of child behaviors rather than the segmentation of clinician actions. The RJA protocol involves detecting relatively simple and overt biomarkers (such as head and gaze shifts toward targets), which current child behavior toolkits can robustly capture. In stark contrast, the ENP protocol requires multimodal analysis of the child, including the simultaneous evaluation of gesture topology (index-finger pointing vs. whole-hand pointing), pointing vector accuracy, and eye contact under occlusion. Children’s behaviors are often fleeting and unpredictable, and they may be partially occluded in video recordings. Current perception models still exhibit limitations in reliably interpreting these subtle and multidimensional actions, leading to higher scoring errors in the ENP protocol.
5.4 Limitations and future work
A key limitation of this study is that the data were collected from a single clinical center involving only 48 children using a standardized hardware platform. This controlled environment may cause the model to overfit to specific lighting conditions, room backgrounds, and camera angles. Furthermore, a sample size of 48 children and the operational styles of a limited number of clinicians may not fully represent the broad behavioral diversity and operational variations present in ASD and typical development populations. Future work needs to validate the framework across multiple centers with more diverse environments and participant demographics.
Additionally, the decoupled analysis architecture itself presents certain limitations by failing to capture the higher-level semantics of dyadic social interactions. The analysis overlooks the real-time feedback loop between the clinician and the child, relying on feature traversal rather than high-level contextual understanding to identify child biomarkers. Although current rules suffice for established protocols, this limitation might constrain the framework’s compatibility with future, more complex, and highly interactive screening paradigms. For such scenarios, exploring temporal action detection (TAD) techniques to directly locate events within observation windows could yield more context-aware analysis.
To address the bottleneck in parsing child behavior, future research should focus on developing high-performance deep learning models tailored specifically for pediatric populations, constructing larger and richer datasets of children’s gestures and gaze patterns to enhance robustness against occlusion and erratic movements. Simultaneously, the modular design of the framework inherently supports multimodal expansion. Future iterations could integrate data streams beyond computer vision, such as incorporating audio cues or physiological signals like infrared thermal imaging. By correlating physiological responses with social stimuli within temporally defined windows, these modalities can help disambiguate visual cues and enable a more accurate interpretation of children’s intent.
6 Conclusions
In this paper, we introduced a novel computer-aided framework for analyzing clinician–child interactions in early ASD screening, with TAS as its core component. By accurately segmenting the clinician’s actions according to clinical protocols, our framework creates precise temporal windows for analyzing a child’s behavioral biomarkers, alleviating the reliance on labor-intensive manual video trimming or imprecise external triggers. Our specialized LAPS-AS model, which combines part-level motion analysis with language-based semantic guidance, demonstrates state-of-the-art performance across two distinct and challenging clinical protocols. Experiments on a real-world dataset of 96 screening sessions show that our assisted pipeline achieves high consistency with scores derived from manual segmentation, validating the effectiveness of action segmentation for this application. While the scoring for the simpler RJA protocol aligns closely with expert assessment, the results for the more complex ENP protocol highlight that future work should focus on improving the robustness of multi-modal child behavior perception. Overall, our work establishes a powerful and scalable foundation for the next generation of objective, decision-support tools for early ASD screening.
American Psychiatric Association. Diagnostic and statistical manual of mental disorders. 5th ed. Washington: American Psychiatric Association Publishing, 2022
[2]
Maenner M J, Warren Z, Williams A R. et al. Prevalence and characteristics of autism spectrum disorder among children aged 8 years—autism and developmental disabilities monitoring network, 11 sites, United States, 2020. MMWR Surveillance Summaries, 2023, 72(2): 1–14
[3]
Zhou H, Xu X, Yan W L. et al. Prevalence of autism spectrum disorder in China: a nationwide multi-center population-based study among children aged 6 to 12 years. Neuroscience Bulletin, 2020, 36(9): 961–971
[4]
Wang Z Y, Liu J J, Zhang W Q. et al. Diagnosis and intervention for children with autism spectrum disorder: a survey. IEEE Transactions on Cognitive and Developmental Systems, 2022, 14(3): 819–832
[5]
Alcañiz Raya M, Chicchi Giglioli I A, Marín-Morales J. et al. Application of supervised machine learning for behavioral biomarkers of autism spectrum disorder based on electrodermal activity and virtual reality. Frontiers in Human Neuroscience, 2020, 14: 90
[6]
Hurwitz S, Yirmiya N. Autism diagnostic observation schedule (ADOS) and its uses in research and practice. In: Patel V B, Preedy V R, Martin C R, eds. Comprehensive Guide to Autism. New York: Springer, 2014, 345–353
[7]
Rutter M, LeCouteur A, Lord C. Autism diagnostic interview-revised. American Journal of Mental Retardation, 2003
[8]
Shamhan A N M, Qaraqe M, Al-Thani D. Advancements in automated assessment and diagnosis of autism spectrum disorder through multimodality sensing technologies: survey of the last decade. IEEE Transactions on Cognitive and Developmental Systems, 2025, 17(4): 727–745
[9]
De Belen R A J, Bednarz T, Sowmya A. et al. Computer vision in autism spectrum disorder research: a systematic review of published studies from 2009 to 2019. Translational Psychiatry, 2020, 10(1): 333
[10]
Dénes-Fazakas L, Mateas I C, Berciu A G. et al. A real time multi modal computer vision framework for automated autism spectrum disorder screening. Electronics, 2026, 15(6): 1287
[11]
Yang Z Q, Zhang Y Y, Ning J C. et al. Early diagnosis of autism: a review of video-based motion analysis and deep learning techniques. IEEE Access, 2025, 13: 2903–2928
[12]
Deng S J, Kosloski E E, Patel S. et al. Hear me, see me, understand me: audio-visual autism behavior recognition. IEEE Transactions on Multimedia, 2025, 27: 2335–2346
[13]
Shi Y H, Ren W H, Jiang W B, et al. Vision-based action detection for RTI protocol of ASD early screening. In: Proceedings of the 15th International Conference on Intelligent Robotics and Applications. Harbin, China, 2022, 370–380
[14]
Hirsch R, Cohen R, Golany T, et al. Random walks for temporal action segmentation with timestamp supervision. In: Proceedings of 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA, 2024, 6600–6610
[15]
Xu J L, Yin S B, Peng Y X. Human-centric fine-grained action quality assessment. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(8): 6242–6255
[16]
Cheng M, Zhang Y Y, Xie Y X. et al. Computer-aided autism spectrum disorder diagnosis with behavior signal processing. IEEE Transactions on Affective Computing, 2023, 14(4): 2982–3000
[17]
Liu J J, Wang Z Y, Xu K. et al. Early screening of autism in toddlers via response-to-instructions protocol. IEEE Transactions on Cybernetics, 2022, 52(5): 3914–3924
[18]
De Giacomo A, Palmieri R, Russo E F. et al. Machine learning and deep learning applied to EEG and fNIRS for early autism spectrum disorder diagnosis: a systematic review. Frontiers in Psychiatry, 2026, 17: 1668914
[19]
Hatim H A, Alyasseri Z A A, Jamil N. A recent advances on autism spectrum disorders in diagnosing based on machine learning and deep learning. Artificial Intelligence Review, 2025, 58(10): 313
[20]
Wang Z Y, Liu J J, He K S. et al. Screening early children with autism spectrum disorder via response-to-name protocol. IEEE Transactions on Industrial Informatics, 2021, 17(1): 587–595
[21]
Qaraqe M, Varghese E B, Qadir I. et al. Joint attention in autism: a narrative review of assessment techniques from behavioral observation to artificial intelligence. Behavior Research Methods, 2026, 58(3): 77
[22]
Zhang K, Chen J Y, Yang Z Y. et al. Investigating joint attention in children with autism spectrum disorder through virtual reality and eye-tracking: a comparative study. Education and Information Technologies, 2025, 30(13): 18779–18798
[23]
Wang Z Y, Xu K, Liu H H. Screening early children with autism spectrum disorder via expressing needs with index finger pointing. In: Proceedings of the 13th International Conference on Distributed Smart Cameras. Trento, Italy, 2019, 24
[24]
Wang Z Y, Qin H B, Liu J J. et al. Early screening of autism in toddlers via express-needs-with-pointing protocol. IEEE Journal of Biomedical and Health Informatics, 2025, 29(4): 2911–2921
[25]
Ding G D, Sener F, Yao A. Temporal action segmentation: an analysis of modern techniques. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(2): 1011–1030
[26]
Wang T Q, Todorovic S. Timestamp query transformer for temporal action segmentation. In: Proceedings of 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Tucson, USA, 2026, 5016–5025
[27]
Zheng Z C, Zhou Y, Chen Y. et al. What, when and where: spatial-aware temporal action segmentation. Pattern Recognition, 2025, 168: 111833
[28]
Zhang J R, Wen W J, Liu S L. et al. End-to-end streaming video temporal action segmentation with reinforcement learning. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(8): 15449–15462
[29]
Lei P, Todorovic S. Temporal deformable residual networks for action segmentation in videos. In: Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA, 2018, 6742–6751
[30]
Yan S J, Xiong Y J, Lin D H. Spatial temporal graph convolutional networks for skeleton-based action recognition. In: Proceedings of the 32nd AAAI Conference on Artificial Intelligence. New Orleans, USA, 2018
[31]
Ghosh P, Yao Y, Davis L S, et al. Stacked spatio-temporal graph convolutional networks for action segmentation. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Snowmass, USA, 2020, 565–574
[32]
Zhang S, Tong H H, Xu J J. et al. Graph convolutional networks: a comprehensive review. Computational Social Networks, 2019, 6(1): 11
[33]
Li Y H, Liu K Y, Liu S L. et al. Involving distinguished temporal graph convolutional networks for skeleton-based temporal action segmentation. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(1): 647–660
[34]
Filtjens B, Vanrumste B, Slaets P. Skeleton-based action segmentation with multi-stage spatial-temporal graph convolutional neural networks. IEEE Transactions on Emerging Topics in Computing, 2024, 12(1): 202–212
[35]
Lea C, Flynn M D, Vidal R, et al. Temporal convolutional networks for action segmentation and detection. In: Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, USA, 2017, 1003–1012
[36]
Lea C, Vidal R, Reiter A, et al. Temporal convolutional networks: a unified approach to action segmentation. In: Proceedings of the European Conference on Computer Vision. Amsterdam, The Netherlands, 2016, 47–54
[37]
Farha Y A, Gall J. MS-TCN: multi-stage temporal convolutional network for action segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach, USA, 2019, 3570–3579
[38]
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA, 2017, 6000–6010
[39]
Ji H Y, Liu X T, Gao Y, et al. LaDy: Lagrangian-dynamic informed network for skeleton-based action segmentation via spatial-temporal modulation. arXiv preprint: arXiv: 2603.24097, 2026
[40]
Ji H Y, Chen B W, Yang Z H, et al. Spectral scalpel: amplifying adjacent action discrepancy via frequency-selective filtering for skeleton-based action segmentation. arXiv preprint: arXiv: 2603.24134, 2026
[41]
Ji H Y, Chen B W, Ren W H. et al. Text-derived relational graph-enhanced network for skeleton-based action segmentation. IEEE Transactions on Image Processing, 2025, 34: 7305–7320
[42]
Ji H Y, Chen B W, Xu X L, et al. Language-assisted skeleton action understanding for skeleton-based temporal action segmentation. In: Proceedings of the 18th European Conference on Computer Vision. Milan, Italy, 2024, 400–417
[43]
Ji H Y, Chen B W, Huang W Z. et al. Snippet-aware transformer with multiple action elements for skeleton-based action segmentation. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(9): 17462–17476
[44]
Yang Z H, Ji H Y, Chen B W. et al. Topology-motion decoupling framework with textual regularization for skeleton-based temporal action segmentation. IEEE Transactions on Circuits and Systems for Video Technology, 2026, 36(5): 7236–7249
[45]
Chen B W, Nie W, Ji H Y. et al. Multiscale skeleton-based temporal action segmentation using hierarchical temporal modeling and prediction ensemble. IEEE Transactions on Cybernetics, 2025, 55(6): 2779–2791
[46]
Chen B W, Ji H Y, Ma H W. et al. Exploring supervised contrastive learning for skeleton-based temporal action segmentation. IEEE Transactions on Cognitive and Developmental Systems, 2025, 17(4): 964–975
[47]
Li Q M, Han Z C, Wu X M. Deeper insights into graph convolutional networks for semi-supervised learning. In: Proceedings of the 32nd AAAI Conference on Artificial Intelligence. New Orleans, USA, 2018
[48]
Chen B W, Ren W H, Liu H H, et al. AutoENP: an auto rating pipeline for expressing needs via pointing protocol. In: Proceedings of 2022 26th International Conference on Pattern Recognition (ICPR). Montreal, Canada, 2022, 3280–3286
[49]
Helvaci H I, Chuah C N, Ozonoff S, et al. Localizing moments of actions in untrimmed videos of infants with autism spectrum disorder. In: Proceedings of 2024 IEEE International Conference on Image Processing (ICIP). Abu Dhabi, United Arab Emirates, 2024, 3841–3847
[50]
Kojovic N, Natraj S, Mohanty S P. et al. Using 2D video-based pose estimation for automated prediction of autism spectrum disorders in young children. Scientific Reports, 2021, 11(1): 15069
[51]
Jabbar U, Waseem Iqbal M, Nechifor A. et al. Deep learning based approach for behavior classification in diagnoses of autism spectrum disorder using naturalistic videos. Frontiers in Computational Neuroscience, 2026, 20: 1626315
[52]
Lyu C Q, Zhang W W, Huang H A, et al. RTMDet: an empirical study of designing real-time object detectors. arXiv preprint: arXiv: 2212.07784, 2022
[53]
He K M, Zhang X Y, Ren S Q, et al. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas, USA, 2016, 770–778
[54]
Zhou Y J, Gregson J. WHENet: real-time fine-grained estimation for wide range head pose. arXiv preprint: arXiv: 2005.10353, 2020
[55]
Zhang X C, Park S, Beeler T, et al. ETH-XGaze: a large scale dataset for gaze estimation under extreme head pose and gaze variation. In: Proceedings of the European Conference on Computer Vision. Glasgow, UK, 2020, 365–381
[56]
Liu S L, Zeng Z Y, Ren T H, et al. Grounding DINO: marrying DINO with grounded pre-training for open-set object detection. In: Proceedings of the 18th European Conference on Computer Vision. Milan, Italy, 2024, 38–55
[57]
Wan S, Gong C, Zhong P. et al. Multiscale dynamic graph convolutional network for hyperspectral image classification. IEEE Transactions on Geoscience and Remote Sensing, 2020, 58(5): 3162–3177
[58]
Yin P Z, Nie J, Liang X Y. et al. A multiscale graph convolutional neural network framework for fault diagnosis of rolling bearing. IEEE Transactions on Instrumentation and Measurement, 2023, 72: 2520713
[59]
Jang S, Lee H, Kim W J. et al. Multi-scale structural graph convolutional network for skeleton-based action recognition. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(8): 7244–7258
[60]
Ishikawa Y, Kasai S, Aoki Y, et al. Alleviating over-segmentation errors by detecting action boundaries. In: Proceedings of 2021 IEEE Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA, 2021, 2321–2330
[61]
Lu Z J, Elhamifar E. FACT: frame-action cross-attention temporal modeling for efficient action segmentation. In: Proceedings of 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA, 2024, 18175–18185
[62]
Li Y H, Li Z Y, Gao S H, et al. A decoupled spatio-temporal framework for skeleton-based action segmentation. arXiv preprint: arXiv: 2312.05830, 2023
Rights & permissions
The Author(s) 2026. This article is published by Higher Education Press.