MuSE‐Promoter: a multi-scale feature fusion and weighted ensemble learning method for identifying promoters across multiple cell lines

Xiao Bi , Zhangyu Mei , Hao Wu

Molecular and Digital Medicine ›› 2026, Vol. 1 ›› Issue (1) : 100002

PDF (6868KB)
Molecular and Digital Medicine ›› 2026, Vol. 1 ›› Issue (1) :100002 DOI: 10.1016/j.mdmed.2026.100002
Original Research
research-article
MuSE‐Promoter: a multi-scale feature fusion and weighted ensemble learning method for identifying promoters across multiple cell lines
Author information +
History +
PDF (6868KB)

Abstract

Promoters are central to regulating gene transcription by orchestrating cell-type- and developmental-stage-specific expression programs. However, the intrinsic heterogeneity of transcription factor binding sites, characterized by variable lengths, complex combinatorial patterns, and substantial sequence diversity across cell types, poses significant challenges to the robustness and generalizability of models relying on single-feature representations. To address these limitations, we propose MuSE-Promoter, a deep ensemble framework that integrates multi-scale feature fusion with weighted ensemble learning for accurate promoter identification across diverse cell lines. MuSE-Promoter constructs parallel feature extraction channels that combine contextual sequence embeddings from the DNABERT model and Word2Vec embeddings with handcrafted descriptors, including tri-nucleotide physicochemical properties (TPCP) and reverse-complement k-mer frequencies (RCKmer). A multi-scale convolutional neural network augmented with squeeze-and-excitation (SE) attention captures hierarchical motif patterns while effectively suppressing noise, followed by a Transformer module to model long-range dependencies. Furthermore, the deep learning branch is integrated with a Random Forest classifier through a learnable weighted ensemble strategy, thereby enhancing cross-domain robustness and prediction stability. Comprehensive evaluations across human cell lines from multiple tissues and Arabidopsis thaliana datasets demonstrate that MuSE-Promoter consistently outperforms state-of-the-art methods. Notably, it achieves superior generalization performance in challenging scenarios, including cross-cell-line transfer and enhancer–promoter discrimination. Collectively, MuSE-Promoter provides a powerful computational framework for large-scale promoter annotation, offering new insights into the regulatory mechanisms underlying complex transcriptional regulation.

Keywords

MuSE-Promoter / Promoter identification / Multi-scale feature fusion / Transformer / Ensemble learning / Cross-domain generalization / DNABERT / Word2Vec / TPCP / RCKmer

Cite this article

Download citation ▾
Xiao Bi, Zhangyu Mei, Hao Wu. MuSE‐Promoter: a multi-scale feature fusion and weighted ensemble learning method for identifying promoters across multiple cell lines. Molecular and Digital Medicine, 2026, 1 (1) : 100002 DOI:10.1016/j.mdmed.2026.100002

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Haberle V, Stark A. Eukaryotic core promoters and the functional basis of transcription initiation. Nat Rev Mol Cell Biol. 2018; 19(10): 621-637.

[2]

Griffith EC, West AE, Greenberg ME. Neuronal enhancers fine—tune adaptive circuit plasticity. Neuron. 2024; 112(18): 3043-3057.

[3]

Zhao X, Song L, Yang A, et al. Prioritizing genes associated with brain disorders by leveraging enhancer—promoter interactions in diverse neural cells and tissues. Genome Med. 2023; 15(1): 56.

[4]

Yang JH, Hansen AS. Enhancer selectivity in space and time: from enhancer—promoter interactions to promoter activation. Nat Rev Mol Cell Biol. 2024; 25(7): 574-591.

[5]

Kawasaki K, Fukaya T. Regulatory landscape of enhancer—mediated transcriptional activation. Trends Cell Biol. 2024; 34(10): 826-837.

[6]

Gulati GS, D’Silva JP, Liu Y, et al. Profiling cell identity and tissue architecture with single—cell and spatial transcriptomics. Nat Rev Mol Cell Biol. 2025; 26(1): 11-31.

[7]

Wang X, Li F, Zhang Y, et al. Deep learning approaches for non—coding genetic variant effect prediction: current progress and future prospects. Brief Bioinforma. 2024; 25(5): bbae446.

[8]

Zhang Y, Zhang P, Wu H. Enhancer—MDLF: a novel deep learning framework for identifying cell—specific enhancers. Brief Bioinforma. 2024; 25(2): bbae083.

[9]

Zhou X, Wu H. scHiClassifier: a deep learning framework for cell type prediction by fusing multiple feature sets from single—cell Hi—C data. Brief Bioinforma. 2024; 26(1): bbaf009.

[10]

Lin H, Deng E—Z, Ding H, et al. iPro54—PseKNC: a sequence—based predictor for identifying sigma—54 promoters in prokaryote with pseudo k—tuple nucleotide composition. Nucleic Acids Res. 2014; 42(21): 12961-12972.

[11]

Lin H, Liang Z—Y, Tang H, Chen W. Identifying sigma70 promoters with novel pseudo nucleotide composition. IEEE/ACM Trans Comput Biol Bioinforma. 2017; 16(4): 1316-1321.

[12]

Liu B, Yang F, Huang D—S, Chou K—C. iPromoter—2L: a two—layer predictor for identifying promoters and their types by multi—window—based PseKNC. Bioinformatics. 2018; 34(1): 33-40.

[13]

Xiao X, Xu Z—C, Qiu W—R, et al. iPSW (2L)—PseKNC: a two—layer predictor for identifying promoters and their strength by hybrid features via pseudo K—tuple nucleotide composition. Genomics. 2019; 111(6): 1785-1793.

[14]

Chevez—Guardado R, Peña—Castillo L. Promotech: a general tool for bacterial promoter recognition. Genome Biol. 2021; 22(1): 318.

[15]

Umarov RK, Solovyev VV. Recognition of prokaryotic and eukaryotic promoters using convolutional deep learning neural networks. PLOS One. 2017; 12(2): e0171410.

[16]

Oubounyt M, Louadi Z, Tayara H, Chong KT. DeePromoter: robust promoter predictor using deep learning. Front Genet. 2019; 10: 286.

[17]

Zhang P, Zhang H, Wu H. iPro—WAEL: a comprehensive and robust framework for identifying promoters in multiple species. Nucleic Acids Res. 2022; 50(18): 10278-10289.

[18]

Rane NL, Paramesha M, Choudhary SP, Rane J. Machine learning and deep learning for big data analytics: a review of methods and applications. Partn Univers Int Innov J. 2024; 2(3): 172-197.

[19]

Abbasi AF, Asim MN, Dengel A. Transitioning from wet lab to artificial intelligence: a systematic review of AI predictors in CRISPR. J Transl Med. 2025; 23(1): 153.

[20]

Barbadilla—Martínez L, Klaassen N, van Steensel B, de Ridder J. Predicting gene expression from DNA sequence using deep learning models. Nat Rev Genet. 2025: 1-15.

[21]

Wu YF, Shi ZQ, Zhou XF, et al. scHiCyclePred: a deep learning framework for predicting cell cycle phases from single—cell Hi—C data using multi—scale interaction information. Commun Biol. 2024; 7(1): 923.

[22]

Yang XH, Mann KK, Wu H, Ding J. scCross: a deep generative model for unifying single—cell multi—omics with seamless integration, cross—modal generation, and in silico exploration. Genome Biol. 2024; 25(1): 198.

[23]

Wang X, Xu K, Huang Z, et al. Accelerating promoter identification and design by deep learning. Trends Biotechnol. 2025; 43(12): 3071-3087.

[24]

Liu H, Guo F, Huang L, et al. Dbert2_LR: a deep learning—based model for predicting cis—regulatory elements in crops. Genomics. 2026; 118(2): 111201.

[25]

Molho D, Ding J, Tang W, et al. Deep learning in single—cell analysis. ACM Trans Intell Syst Technol. 2024; 15(3): 1-62.

[26]

Sherman MA. Genetics Disease Algorithms Decode Somatic Mutations. Massachusetts Institute Technology; 2023.

[27]

Xue S, Zhang Z, Yan F, et al. WHANet: wavelet and hybrid attention network for vessel segmentation in OCTA fundus images. J Supercomput. 2025; 81(14): 1333.

[28]

Zeng L, Li Z, Shen E, et al. PAIRNet: Predicting PIWI cleavage specificity via position—aware RNA interaction modeling. PLOS Comput Biol. 2026; 22(2): e1013936.

[29]

Yang T, Henao R. TAMC: a deep—learning approach to predict motif—centric transcriptional factor binding activity based on ATAC—seq profile. PLOS Comput Biol. 2022; 18(9): e1009921.

[30]

Zhai J, Zhang Y, Zhang C, et al. deepTFBS: improving within‐and cross‐species prediction of transcription factor binding using deep multi‐task and transfer learning. Adv Sci. 2025; 12(30): e03135.

[31]

Chen J—A, Niu W, Ren B, et al. Survey: exploiting data redundancy for optimization of deep learning. ACM Comput Surv. 2023; 55(10): 1-38.

[32]

Khan A, Rauf Z, Sohail A, et al. A survey of the vision transformers and their CNN—transformer based variants. Artif Intell Rev. 2023; 56(3): 2917-2970.

[33]

Li Y, Cheng Z, Zhang Y, Xu R. QPCNet: a hybrid quantum positional encoding and channel attention network for image classification. Phys Scr. 2025; 100(11): 115101.

[34]

Huminiecki Ł, Horbańczuk J. Can we predict gene expression by understanding proximal promoter architecture? Trends Biotechnol. 2017; 35(6): 530-546.

[35]

Zhang C—T, Zhang R, Ou H—Y. The Z curve database: a graphic representation of genome sequences. Bioinformatics. 2003; 19(5): 593-599.

[36]

Bouïs D, Hospers GA, Meijer C, et al. Endothelium in vitro: a review of human vascular endothelial cell lines for blood vessel—related research. Angiogenesis. 2001; 4(2): 91-102.

[37]

Sajan SA, Hawkins RD. Methods for identifying higher—order chromatin structure. Annu Rev Genom Hum Genet. 2012; 13(1): 59-82.

PDF (6868KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/