Introduction
An AI-driven research paradigm shift
We are now on the eve of a paradigm shift. This shift is moving us into what is often called “the Fourth paradigm of science”, a data-driven paradigm[
1]. This new paradigm does not replace the traditional, hypothesis-driven paradigm. Instead, it reorders the scientific process: it prioritizes the large-scale, data-driven discovery over hypothesis-driven research. This transformation is born not from a single breakthrough, but from the unprecedented convergence of three powerful technological waves: scalable cloud infrastructure, advanced AI architectures—represented by the Transformer[
2]—for complex pattern discovery, and initiatives for large-scale biological data standardization to ensure signal consistency. Together, these waves provide the foundation for AI-powered tools that can uncover insights at a previously unimaginable scale. However, this wave of innovation has yet to breach the formidable barrier in medical research: the neurodegenerative diseases (NDDs). This barrier is solidified by the “silent phase”—the temporal disconnection between the silent progression of molecular pathology and the eventual appearance of clinical symptoms, particularly within the research of Alzheimer’s disease (AD)[
3]. The combination of its long latent period and frequent neuropathological comorbidity turns AD from a single disease trajectory into a complex spatiotemporal network. Recent evidence from a Chinese community-based autopsy cohort further supports this mixed-pathology reality[
4]. Here lies a paradox: the data-driven paradigm is a core driving paradigm to unlock the spatiotemporal complexity of AD, yet it remains starved of input. This data void creates a translational bottleneck. To empower the paradigm to finally breach the barrier, we require a new source of fuel—a pre-symptomatic biomarker that is scalable and readily detectable[
5].
This critical demand has spurred the revival of electroencephalography (EEG). The NIA-AA ATN framework has redefined AD around Amyloid, Tau, and Neurodegeneration, providing a biological anchor beyond symptom-based disease[
6]. Yet this framework leaves the functional dynamics of AD insufficiently captured. While historically viewed as a classical tool, EEG is uniquely positioned to address this gap[
7]. Advances in modern machine learning have substantially expanded its analytical potential, transforming it from a simple chart into a high-dimensional data stream[
8,
9]. Its characteristics are perfectly suited to the new data-driven paradigm: unlike metabolic or structural neuroimaging (PET/MRI), it offers millisecond-level temporal resolution to capture real-time neural dynamics; unlike invasive fluid biopsies, it is non-invasive and cost-effective[
10]. However, EEG also has clear limitations, including low spatial resolution, weak sensitivity to deep neural sources, and vulnerability to artefacts. Therefore, EEG should not be viewed as a replacement for molecular, imaging or neuropathological markers, but as a functional window onto network-level dysfunction. Thus, EEG serves not merely as a measurement tool, but as a critical bridge connecting the computational discovery with clinical-biological validation. For AD, this bridge is particularly important. Data-driven models can search for latent electrophysiological signatures across the silent and prodromal phases, while hypothesis-driven research remains essential for testing their clinical and biological validity against cognitive trajectories, established biomarkers, neuroimaging findings and post-mortem pathology. However, a significant barrier prevents the full realisation of this potential. At present, EEG data remain trapped in fragmented and heterogeneous silos. Without comprehensive standardization, spanning from acquisition protocols, data formats, clinical metadata to privacy governance, raw recordings cannot be transformed into trustworthy, clinically validatable evidence.
This perspective review aims to provide a blueprint for breaking the current deadlock. We first map the landscape of EEG databases, ranging from foundational archives in sleep, epilepsy, and emotion research to emerging AD collections. Their value is immense. However, they remain largely isolated and cannot be readily integrated due to the absence of a common standard. We argue that the field requires a dual strategy: prospective standardization to prevent further irreversible information loss in the future EEG cohorts, and retrospective harmonisation to maximise the value of existing heterogeneous datasets. This is not merely a regulatory formality, but a technical and biological necessity for reproducibility, interoperability and clinical validation. It is driven by the critical need to transform technical model outputs into trustworthy, clinically validatable evidence. Finally, drawing on direct experience with the comprehensive Standardized Operational Protocol (SOP) for the China Human Brain Bank Consortium (CHBBC)[
11,
12], we propose a concrete initiative to build this future: a framework for shared data to accelerate collaborative discovery in NDDs, particularly in AD.
The EEG database landscape: the potential in fragments
Our blueprint for the AI-driven AD research begins with the raw material: data. The field is not starting from nothing. Pioneering efforts have yielded vast repositories, creating an apparent wealth of information. This chapter maps this complex terrain.
Our survey begins with the representative foundational EEG archives and platforms established in classic domains, including sleep, epilepsy, and emotion research (summarized in Table 1). These resources demonstrate both the feasibility of large-scale data collection and the practical diversity of existing sharing models, thereby providing valuable experience for developing future databases.
As Table 1 illustrates, repositories such as Temple University Hospital (TUH) EEG Corpus have achieved an unprecedented scale, providing the volume necessary for developing robust deep learning algorithms. Concurrently, platforms such as PhysioNet and OpenNeuro have been instrumental in fostering a culture of open data sharing.
Many of these archives are well curated, widely cited, and have supported decades of productive research. However, they were constructed for domain-specific research questions rather than for cross-disease integration or large-scale translational application. Table 1 highlights several dimensions of heterogeneity that directly affect cross-dataset integration:
• Sampling rate: explicitly reported sampling rates range from 100 Hz to 1,000 Hz, while several repositories further depend on dataset- or device-specific settings.
• Channel configuration: channel counts vary from low-density or device-dependent recordings to 118-channel configurations, and several archives do not enforce a uniform montage.
• File format: data formats include .edf, .bdf, .cnt, .mat, .txt, .pkl, and BIDS-compatible structures, creating substantial preprocessing barriers before joint analysis.
• Accessibility and reusability: beyond technical heterogeneity, several resources are affected by unstable links, incomplete documentation, or platform-dependent access restrictions, further limiting practical reuse.
These differences mean that direct integration of otherwise valuable datasets can introduce acquisition-, format-, and access-related confounders that compromise scientific rigour. The field is therefore left in a paradoxical situation: possessing a wealth of data that it cannot collectively mine.
The experience with foundational databases exposes a universal challenge. The issue of data fragmentation becomes even more acute when we turn our attention to the specific and urgent challenge of AD.
Unlike in foundational domains such as sleep, epilepsy, and emotion research, the large-scale, standardized public EEG databases specifically designed for AD are exceedingly rare. The majority of studies, typically involving only dozens of participants, do not release their raw data to the public. The few studies that have made their data available are summarized in Table 2. Consequently, the data landscape for AD is even more challenging than for foundational research: available datasets are not only highly heterogeneous but also severely limited in quantity.
The evidence in Table 2 solidifies this conclusion. To make this imbalance explicit, Figure 1 compares the participant numbers of publicly available AD-focused EEG datasets with the TUHEEG Corpus as a foundational benchmark. The comparison shows that most public AD-focused EEG datasets remain at the scale of dozens to hundreds of participants. Taken together, the foundational archives (Table 1) and the emerging AD datasets (Table 2) form a fragmented landscape of potential. This landscape of valuable but disconnected data is fundamentally incapable of powering the large-scale, data-driven research we envision. This is the central bottleneck. The solution to this paralysis is standardization. In the following chapter, we will propose a concrete framework to implement it.
Standardization: the cornerstone of the new paradigm
The fragmentation of existing data does not merely cause inconvenience; it leads research astray. Integrating heterogeneous data, particularly by merging EEG records with different sampling rates, filter settings, or electrode montages, is scientifically unsound. It inevitably introduces profound biases that can obscure true biological signals or generate spurious findings[
13–
15]. While modern large model architectures demonstrate increasing technical resilience to heterogeneous inputs, relying on such inconsistency for clinical discovery remains a fundamental risk. Establishing robust, clinically reproducible biomarkers demands a harmonised baseline, regardless of the model’s capacity. Feeding these models a patchwork of inconsistent data guarantees “garbage in, garbage out” rendering any resulting biomarker unreliable for medical applications[
13,
16]. Furthermore, this lack of a common standard fuels the reproducibility crisis in neuroscience, including NDD research[
16], and impedes the development of automated, high-throughput analysis pipelines, trapping researchers in a Sisyphean cycle of writing bespoke scripts[
17].
Here, we distinguish between prospective standardization and retrospective harmonisation. Prospective standardization refers to predefined protocols applied before data generation, including clinical metadata, acquisition parameters, ethical governance, and file organisation. Its purpose is to prevent irreversible information loss. Retrospective harmonization, by contrast, refers to post-acquisition strategies that improve the interoperability of already collected heterogeneous datasets, including BIDS conversion, signal normalisation, metadata mapping, domain adaptation, and, federated learning. These approaches complement, rather than replace, one another.
For future AD-oriented EEG cohorts, the core instrument of prospective standardization is a comprehensive SOP. To address the full spectrum of challenges from compliance to analysis, this SOP should include four levels of content:
(1) Ethical and procedural standardization: establishing strict protocols for data desensitisation, privacy protection, consent for longitudinal follow-up, and compliance governance to ensure the legality of large-scale, multi-centre sharing;
(2) Clinical standardization: using Common Data Elements (CDEs)[
18] for demographic information, cognitive scales, diagnostic criteria, medication status, comorbidities, biomarker status, and longitudinal outcomes;
(3) Acquisition standardization: defining unified sampling rates, filters, reference schemes, electrode montages, impedance criteria, recording duration, and task status, such as eyes-closed resting-state EEG and selected task-related paradigms, to maximise the retention of original information while minimising irreversible operations;
(4) Format standardization: mandating the community-wide adoption of the Brain Imaging Data Structure (BIDS)[
19,
20], which transforms isolated files into a unified, machine-readable, analysis-ready dataset. For newly collected EEG, BIDS should be embedded prospectively into the SOP; for legacy datasets, BIDS conversion provides a practical route for retrospective harmonisation.
This call for a comprehensive SOP is not merely theoretical. A powerful precedent already exists. The CHBBC has successfully implemented a multi-centre SOP for the complex logistics of human tissue collection[
11,
12,
21,
22]. Crucially, this infrastructure is designed to manage dual repositories: physical biological specimens and their corresponding digital assets (e.g., clinical records, digitised pathology, and omics). Therefore, incorporating digital EEG into this archive is not a reconstruction but a logical and seamless extension. Rather than claiming a completed resource, we propose a staged initiative to link ante-mortem functional EEG with post-mortem neuropathology. Longitudinal brain-donation cohorts have already demonstrated the feasibility of connecting ante-mortem clinical data with neuropathological endpoints, including BDR in the UK[
23], ROSMAP in the United States[
24], and BANC-PUMC cohort in China[
25]. Importantly, isolated precedents such as Arizona Study of Aging and Neurodegenerative Disorders and Brain and Body Donation Program (AZSAND/BBDP) have further shown that premortem resting-state EEG can be analyzed in autopsy-confirmed NDD cohorts[
26]. However, the systematic incorporation of standardized EEG into such living-to-post-mortem frameworks remains rare and underdeveloped.
In short, prospective standardization provides the trustworthy baseline for future EEG cohorts, while retrospective harmonization maximises the usability of existing heterogeneous datasets. Together, they build the foundation that the new paradigm scientifically demands.
Analytical evolution: from manual feature extraction to AI-driven discovery
The standardized, large-scale data infrastructure established in the previous chapter provides the essential “clean fuel”. However, this fuel requires more powerful analytical engines to convert it into discovery. This chapter charts the evolution of these engines, transitioning from traditional methods to data-driven approaches enabled by this new paradigm.
Historically, EEG analysis depended on visual inspection and quantitative methods such as spectral analysis. These approaches were foundational, translating the raw signals into interpretable features based on prior hypotheses. This process, while essential for interpretability, is limited by its design, risking the oversight of complex, non-linear patterns[
27]. The field evolved towards more advanced feature engineering, using techniques such as functional connectivity and microstate analysis. This knowledge-driven approach peaked in systems such as the dementia screen proposed by Huang et al.[
28]. By manually crafting novel non-linear features, they achieved remarkable classification accuracy (AUC > 0.91). This success proves the immense value hidden in EEG signals, but it also reveals the fundamental limitation: the discovery of features remains a manual, hypothesis-bound process.
This leads to end-to-end (E2E) deep learning. A common view frames this as a complete replacement for manual feature extraction. This perspective is inaccurate. E2E learning does not bypass human knowledge; it abstracts it. Our decades of domain expertise are not discarded. Instead, they are encoded into the model’s architecture as inductive biases[
29]. While the Transformer architecture currently spearheads this revolution, it operates within a broader architectural constellation, where each design mirrors a classical neuroscientific hypothesis. Just as Convolutional Neural Networks (CNNs) enforce local receptive fields to capture time-frequency patterns, and Graph Neural Networks (GNNs) model the brain’s topology to mimic functional connectivity, the Transformer’s attention mechanisms[
2] is a direct implementation of the hypothesis that long-range temporal dependencies are critical in EEG signals[
30]. Thus, the E2E model is not an unbiased “black box” but a powerful tool that scales up our human-defined priors.
However, this power is bound by an inescapable reality regarding data. It is true that modern foundation models demonstrate the engineering resilience to ingest heterogeneous inputs—handling inconsistent length or sampling rates via technical adaptations. However, clinical discovery demands more than engineering feasibility; it demands biological validity. Relying on inconsistent data risks training models that learn the “noise” of the acquisition protocol rather than the “signal” of the pathology. Therefore, simply feeding data—following the “Scaling Laws”[
31] —is insufficient if the data source acts as a confounder; rigorous harmonisation of processing pipelines remains a non-negotiable imperative. This data dependency reinforces the central argument of this review: the analytical methods have evolved to a point where our fragmented data landscape (chapter 2) is the primary bottleneck.
To deploy these next-generation discovery engines, we must first build the foundation from chapter 3. However, this synergy between data and models is not automatic, and significant challenges remain. First, robustness is not guaranteed. A model trained on one database can learn site-specific artefacts, especially in the low signal-to-noise ratio EEG data, risking “Shortcut Learning” that compromises clinical validity across centres[
32]. Besides, the issue of interpretability is paramount[
33]. An E2E model may discover predictive patterns that are physiologically opaque, requiring a return to rigorous, hypothesis-driven methods for validation[
34]. These challenges—robustness and interpretability—cannot be solved by a better algorithm alone. Furthermore, they highlight a third deficit: the lack of scalable infrastructure to validate and deploy these models across diverse centres. These systemic problems demand a new ecosystem for collaborative discovery. Building this predictable and reliable blueprint is the task we turn to next.
Future blueprint: a new ecosystem for brain research
The transition from manual analysis to AI-driven discovery faces a “triad of deficits”: Robustness, Interpretability, and Scalability (RIS). As outlined in chapter 4, these are not merely technical footnotes but systemic barriers. They cannot be overcome by isolated algorithmic improvements; they demand a new infrastructure for analysis and a new ecosystem for collaboration. This final chapter proposes a blueprint for this future, built directly upon the three technological waves we introduced at the outset.
The foundation of this solution is the technical infrastructure, and this is where scalable cloud computing becomes the engine of discovery. We envision an open, modular platform that moves beyond simple data storage[
35]. This vision is grounded in reality; successful precedents such as the CHBBC’s existing portal have already demonstrated the feasibility of centralizing resources on a national scale. It is precisely this unified architecture that addresses the first challenge: robustness (R). It allows for large-scale, federated model training and validation across diverse, multi-centre datasets, ensuring that biomarkers are generalizable and not mere site-specific artefacts[
36]. By requiring models to be trained, stress-tested, and externally validated across centres, the platform turns multi-centre heterogeneity from an uncontrolled confounder into an explicit benchmark for robustness.
This cloud platform is also the stage for our new analytical engines. Here, the Transformer architecture (chapter 4) and prospective standardization (chapter 3) converge. The platform hosts advanced E2E models, including Transformer-based EEG models, while the standardized data provides the “clean fuel” they require[
37]. This synergy directly addresses the second challenge: interpretability (I). For EEG, practical explainable artificial intelligence (XAI) should be organized around neurophysiological dimensions rather than generic model explanations alone, using methods such as saliency analysis, occlusion testing, or attention/relevance visualization[
38,
39]. First, temporal attribution can identify which recording windows contribute most to a prediction. Second, spectral attribution can test whether the model relies on biologically plausible frequency components, such as slowing-related change in AD. Third, spatial or connectivity-based attribution can examine whether relevant channels, regions, or network patterns are consistent with known neurodegenerative involvement. Embedding these XAI outputs into the cloud platform would allow different centres to compare, reproduce, and challenge the same model explanations. By validating these explanatory outputs against standardized data, cognitive trajectories, neuroimaging markers, and post-mortem pathology, we can democratize access to advanced computation and collectively translate opaque patterns into physiological understanding.
With this infrastructure and its powerful analytics in place, the platform addresses the final challenge: scalability (S). This is not just about data volume, but about longitudinal depth, and biological dimensionality. This platform becomes the hub for a new collaborative ecosystem. Its true power is unlocked when it integrates data far beyond resting-state EEG. The highest priority is to create synergy with Human Brain Banks, allowing ante-mortem functional dynamics to be correlated with post-mortem neuropathology. Equally important is integration with large-scale prospective cohorts. Retrospective harmonization can maximize the value of existing EEG datasets, but it cannot create the missing silent-phase observations by itself. These observations require prospectively standardized EEG embedded into longitudinal studies that capture cognition, imaging, genomics, and multi-omics. Such cohorts would allow us to track the transition from silent pathology to clinical decline and explore the complex “gene-environment-brain” interactions. This ecosystem must also capture real-world data through wearable devices, moving from clinical snapshots to the longitudinal monitoring required for true personalisation.
This technical and scientific blueprint, however, cannot succeed without a corresponding cultural shift. The technical capacity for scalability is useless without the communal will to share. We must champion a move towards Open Science, grounded in the FAIR Principles (Findable, Accessible, Interoperable, and Reusable). This openness must be built on a foundation of trust. To sustain this ecosystem, the privacy of participants is non-negotiable, requiring state-of-the-art techniques like differential privacy and strong ethical governance. Only by securing trust can we ensure the data flow necessary for a truly robust and scalable science.
The paradigm shift we described in chapter 1 is not a spectator sport. It requires building. The blueprint outlined here—a synergistic platform designed to ensure robustness, clarify interpretability, and enable scalability (RIS)—is the necessary next step. It is not merely a technical roadmap for a database; it is a call for a new, collaborative method of scientific discovery. Although AD serves as the index disease in this perspective, mixed pathology makes the proposed framework relevant to NDDs more broadly. By solving the RIS challenges, we can move from fragmented EEG records to clinically and biologically grounded understanding across the ageing brain.