The gap-free genome of mulberry elucidates the architecture and evolution of polycentric chromosomes

Bi Ma , Honghong Wang , Jingchun Liu , Lin Chen , Xiaoyu Xia , Wuqi Wei , Zhen Yang , Jianglian Yuan , Yiwei Luo , Ningjia He

Horticulture Research ›› 2023, Vol. 10 ›› Issue (7) : 111

PDF (1527KB)
Horticulture Research ›› 2023, Vol. 10 ›› Issue (7) :111 DOI: 10.1093/hr/uhad111
Article
research-article
The gap-free genome of mulberry elucidates the architecture and evolution of polycentric chromosomes
Author information +
History +
PDF (1527KB)

Abstract

Mulberry is a fundamental component of the global sericulture industry, and its positive impact on our health and the environment cannot be overstated. However, the mulberry reference genomes reported previously remained unassembled or unplaced sequences. Here, we report the assembly and analysis of the telomere-to-telomere gap-free reference genome of the mulberry species, Morus notabilis, which has emerged as an important reference in mulberry gene function research and genetic improvement. The mulberry gap-free reference genome produced here provides an unprecedented opportunity for us to study the structure and function of centromeres. Our results revealed that all mulberry centromeric regions share conserved centromeric satellite repeats with different copies. Strikingly, we found that M. notabilis is a species with polycentric chromosomes and the only reported polycentric chromosome species up to now. We propose a compelling model that explains the formation mechanism of new centromeres and addresses the unsolved scientific question of the chromosome fusion-fission cycle in mulberry species. Our study sheds light on the functional genomics, chromosome evolution, and genetic improvement of mulberry species.

Cite this article

Download citation ▾
Bi Ma, Honghong Wang, Jingchun Liu, Lin Chen, Xiaoyu Xia, Wuqi Wei, Zhen Yang, Jianglian Yuan, Yiwei Luo, Ningjia He. The gap-free genome of mulberry elucidates the architecture and evolution of polycentric chromosomes. Horticulture Research, 2023, 10 (7) : 111 DOI:10.1093/hr/uhad111

登录浏览全文

4963

注册一个新账户 忘记密码

Acknowledgments

We thank André Marques (Max Planck Institute for Plant Breeding Research) for his valuable comments on ChIP-seq analysis. We thank Jianming Zeng (University of Macau), and all the members of his bioinformatics team, biotrainee, for generously sharing their experience and codes. We thank Yangqin Xie for his suggestions. We thank Tian Li for helping to deposit the data. We also thank all members of our group.

This project was supported by the National Natural Science Foundation of China (32101544), and the Chongqing Research Pro-gram of Basic Research and Frontier Technology (cstc2021yszx-jcyj0004).

Author contributions

Conceptualization, B.M. and N.H.; funding and resources, B.M., N.H; data production, formal analyses, investigation, and visualization, B.M.; experimentation, H.W., J.L., L.C., X.X.; sample preparation, B.M., H.W., J.L., X.X., W.W., Z.Y., J.Y., Y.L; writing, B.M.; review and editing: N.H. All authors read and approved the final manuscript.

Data availability

All the raw sequencing data and genome assembly generated for this project are deposited at Nation Genomics Data Center under BioProject no. PRJCA015883, and the genome assembly and annotations are also deposited at MorusDB (https://morus.swu.edu.cn). All the materials in this study are available upon request.

Conflict of interest statement

No conflicts of interest declared.

Supplementary data

Supplementary data is available at, Horticulture Research online.

References

[1]

He N, Zhang C, Qi X et al. Draft genome sequence of the mulberry tree Morus notabilis. Nat Commun. 2013; 4: 2445.

[2]

Ma B, Luo Y, Jia L et al. Genome-wide identification and expression analyses of cytochrome P450 genes in mulberry (Morus notabilis). J Integr Plant Biol. 2014; 56: 887-901.

[3]

Li H, Yang Z, Zeng Q et al. Abnormal expression of bHLH3 disrupts a flavonoid homeostasis network, causing differences in pigment composition among mulberry fruits. Hortic Res. 2020; 7: 83.

[4]

Ma B, Xin Y, Kuang L et al. Distribution and characteristics of transposable elements in the mulberry genome. Plant Genome. 2019; 12: 180094.

[5]

Xuan Y, Ma B, Li D et al. Chromosome restructuring and number change during the evolution of Morus notabilis and Morus alba. Hortic Res. 2022; 9: uhab030.

[6]

Li D, Ma B, Xu X et al. MMHub, a database for the mulberry metabolome. Database-Oxford. 2020; 2020: baaa011.

[7]

Xia Z, Dai X, Fan W et al. Chromosome-level genomes reveal the genetic basis of descending Dysploidy and sex determination in Morus plants. Genomics Proteomics Bioinformatics. 2022; 20: 1119-37.

[8]

Jiao F, Luo RS, Dai XL et al. Chromosome-level reference genome and population genomic analysis provide insights into the evolution and improvement of domesticated mulberry (Morus alba). Mol Plant. 2020; 13: 1001-12.

[9]

Jain M, Bansal J, Rajkumar MS et al. Draft genome sequence of Indian mulberry (Morus indica) provides a resource for functional and translational genomics . Genomics. 2022; 114: 110346.

[10]

Li K, Jiang W, Hui Y et al. Gapless indica rice genome reveals synergistic contributions of active transposable elements and segmental duplications to rice genome evolution. Mol Plant. 2021; 14: 1745-56.

[11]

Song JM, Xie WZ, Wang S et al. Two gap-free reference genomes and a global view of the centromere architecture in rice. Mol Plant. 2021; 14: 1757-67.

[12]

Zhang Y, Fu J, Wang K et al. The telomere-to-telomere gap-free genome of four rice parents reveals SV and PAV patterns in hybrid rice breeding. Plant Biotechnol J. 2022; 20: 1642-4.

[13]

Hou X, Wang D, Cheng Z et al. A near-complete assembly of an Arabidopsis thaliana genome . Mol Plant. 2022; 15: 1247-50.

[14]

Naish M, Alonge M, Wlodzimierz P et al. The genetic and epigenetic landscape of the Arabidopsis centromeres. Science. 2021; 374: eabi7489.

[15]

Wang B, Yang X, Jia Y et al. High-quality Arabidopsis thaliana genome assembly with Nanopore and HiFi long reads . Genomics Proteomics Bioinformatics. 2022; 20: 4-13.

[16]

Belser C, Baurens FC, Noel B et al. Telomere-to-telomere gapless chromosomes of banana using nanopore sequencing. Commun Biol. 2021; 4: 1047.

[17]

Deng Y, Liu S, Zhang Y et al. A telomere-to-telomere gap-free reference genome of watermelon and its mutation library provide important resources for gene discovery and breeding. Mol Plant. 2022; 15: 1268-84.

[18]

Navratilova P, Toegelova H, Tulpova Z et al. Prospects of telomere-to-telomere assembly in barley: analysis of sequence gaps in the MorexV3 reference genome. Plant Biotechnol J. 2022; 20: 1373-86.

[19]

Fu A, Zheng Y, Guo J et al. Telomere-to-telomere genome assembly of bitter melon (Momordica charantia L. var. abbreviata Ser.) reveals fruit development, composition and ripening genetic characteristics . Hortic Res. 2023; 10: uhac228.

[20]

Zhang L, Liang J, Chen H et al. A near-complete genome assembly of Brassica rapa provides new insights into the evolution of centromeres . Plant Biotechnol J. 2023; 21: 1022-32.

[21]

Li F, Xu S, Xiao Z et al. Gap-free genome assembly and comparative analysis reveal the evolution and anthocyanin accumulation mechanism of Rhodomyrtus tomentosa. Hortic Res. 2023; 10: uhad005.

[22]

Zhou Y, Xiong J, Shu Z et al. The telomere-to-telomere genome of Fragaria vesca reveals the genomic evolution of Fragaria and the origin of cultivated octoploid strawberry . Hortic Res. 2023; 10: uhad027.

[23]

Tikader A, Kamble CK . Mulberry wild species in India and their use in crop improvement - a review. Aust J Crop Sci. 2008; 2: 64-72.

[24]

Muller H, Gil J, Drinnenberg IA . The impact of centromeres on spatial genome architecture. Trends Genet. 2019; 35: 565-78.

[25]

Cheng H, Concepcion GT, Feng X et al. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021; 18: 170-5.

[26]

Hoencamp C, Dudchenko O, Elbatsh AMO et al. 3D genomics across the tree of life reveals condensin II as a determinant of architecture type. Science. 2021; 372: 984-9.

[27]

Borthakur D, Busov V, Cao XH et al. Current status and trends in forest genomics. Forestry Research. 2022; 2: 11.

[28]

Kille B, Balaji A, Sedlazeck FJ et al. Multiple genome alignment in the telomere-to-telomere assembly era. Genome Biol. 2022; 23: 182.

[29]

Richards EJ, Ausubel FM . Isolation of a higher eukaryotic telomere from Arabidopsis thaliana. Cell. 1988; 53: 127-36.

[30]

Shi X, Cao S, Wang X et al. The complete reference genome for grapevine (Vitis vinifera L.) genetics and breeding . Hortic Res. 2023; 10: uhad061.

[31]

Nie S, Zhao SW, Shi TL et al. Gapless genome assembly of azalea and multi-omics investigation into divergence between two species with distinct flower color. Hortic Res. 2023; 10: uhac241.

[32]

Puizina J, Weiss-Schneeweiss H, Pedrosa-Harand A et al. Karyotype analysis in Hyacinthella dalmatica (Hyacinthaceae) reveals vertebrate-type telomere repeats at the chromosome ends . Genome. 2003; 46: 1070-6.

[33]

Weiss H, Scherthan H . Aloe spp.-plants with vertebrate-like telomeric sequences. Chromosome Res. 2002; 10: 155-64.

[34]

Hofstatter PG, Thangavel G, Lux T et al. Repeat-based holocentromeres influence genome architecture and karyotype evolution. Cell. 2022; 185: 3153-3168.e18.

[35]

Macas J, Avila Robledillo L, Kreplak J et al. Assembly of the 81.6 Mb centromere of pea chromosome 6 elucidates the structure and evolution of metapolycentric chromosomes. PLoS Genet. 2023; 19: e1010633.

[36]

Neumann P, Navratilova A, Schroeder-Reiter E et al. Stretching the rules: monocentric chromosomes with multiple centromere domains. PLoS Genet. 2012; 8: e1002777.

[37]

Xue C, Liu G, Sun S et al. De novo centromere formation in pericentromeric region of rice chromosome 8. Plant J. 2022; 111: 859-71.

[38]

Zhai Z, Wang X, Ding M . Cell Biology. 3rd ed. Beijing, China: Higher Education Press; 2007.

[39]

Ranallo-Benavidez TR, Jaron KS, Schatz MC . GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat Commun. 2020; 11: 1432.

[40]

Marcais G, Kingsford C . A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 2011; 27: 764-70.

[41]

Chen Y, Nie F, Xie SQ et al. Efficient assembly of nanopore reads via highly accurate and intact error correction. Nat Commun. 2021; 12: 60.

[42]

Vaser R, Sovic I, Nagarajan N et al. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 2017; 27: 737-46.

[43]

Walker BJ, Abeel T, Shea T et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One. 2014; 9: e112963.

[44]

Chen S, Zhou Y, Chen Y et al. Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018; 34: i884-90.

[45]

Zhang X, Zhang S, Zhao Q et al. Assembly of allele-aware, chromosomal-scale autopolyploid genomes based on hi-C data. Nat Plants. 2019; 5: 833-45.

[46]

Durand NC, Robinson JT, Shamim MS et al. Juicebox provides a visualization system for hi-C contact maps with unlimited zoom. Cell Syst. 2016; 3: 99-101.

[47]

Jain C, Rhie A, Hansen NF et al. Long-read mapping to repetitive reference sequences using Winnowmap2. Nat Methods. 2022; 19: 705-10.

[48]

Wolff J, Rabbani L, Gilsbach R et al. Galaxy HiCExplorer 3: a web server for reproducible hi-C, capture hi-C and single-cell hi-C data analysis, quality control and visualization. Nucleic Acids Res. 2020; 48: W177-84.

[49]

Manni M, Berkeley MR, Seppey M et al. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 2021; 38: 4647-54.

[50]

Li H, Durbin R . Fast and accurate short read alignment with burrows-wheeler transform. Bioinformatics. 2009; 25: 1754-60.

[51]

Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018; 34: 3094-100.

[52]

Ou S, Chen J, Jiang N . Assessing genome assembly quality using the LTR assembly index (LAI). Nucleic Acids Res. 2018; 46: e126.

[53]

Flynn JM, Hubley R, Goubert C et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc Natl Acad Sci U S A. 2020; 117: 9451-7.

[54]

Ou S, Su W, Liao Y et al. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol. 2019; 20: 275.

[55]

Stanke M, Diekhans M, Baertsch R et al. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics. 2008; 24: 637-44.

[56]

Delcher AL, Bratke KA, Powers EC et al. Identifying bacterial genes and endosymbiont DNA with glimmer. Bioinformatics. 2007; 23: 673-9.

[57]

Kim D, Paggi JM, Park C et al. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019; 37: 907-15.

[58]

Kovaka S, Zimin AV, Pertea GM et al. Transcriptome assembly from long-read RNA-seq alignments with StringTie2. Genome Biol. 2019; 20: 278.

[59]

Holt C, Yandell M . MAKER2: an annotation pipeline and genome-database management tool for second-generation genome projects. BMC Bioinformatics. 2011; 12: 491.

[60]

Kanehisa M, Araki M, Goto S et al. KEGG for linking genomes to life and the environment. Nucleic Acids Res. 2008; 36: D480-4.

[61]

Jones P, Binns D, Chang HY et al. InterProScan 5: genome-scale protein function classification. Bioinformatics. 2014; 30: 1236-40.

[62]

Boeckmann B, Bairoch A, Apweiler R et al. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Res. 2003; 31: 365-70.

[63]

Finn RD, Bateman A, Clements J et al. Pfam: the protein families database. Nucleic Acids Res. 2014; 42: D222-30.

[64]

Ashburner M, Ball CA, Blake JA et al. Gene ontology: tool for the unification of biology. Nat Genet. 2000; 25: 25-9.

[65]

Chan PP, Lin BY, Mak AJ et al. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res. 2021; 49: 9077-96.

[66]

Lagesen K, Hallin P, Rodland EA et al. RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucleic Acids Res. 2007; 35: 3100-8.

[67]

Nawrocki EP, Eddy SR . Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics. 2013; 29: 2933-5.

[68]

Marcais G, Delcher AL, Phillippy AM et al. MUMmer4: a fast and versatile genome alignment system. PLoS Comput Biol. 2018; 14: e1005944.

[69]

Goel M, Sun H, Jiao WB et al. SyRI: finding genomic rearrangements and local sequence differences from whole-genome assemblies. Genome Biol. 2019; 20: 277.

[70]

Zhou ZW, Yu ZG, Huang XM et al. GenomeSyn: a bioinformatics tool for visualizing genome synteny and structural variations. J Genet Genomics. 2022; 49: 1174-6.

[71]

Tang H, Bowers JE, Wang X et al. Synteny and collinearity in plant genomes. Science. 2008; 320: 486-8.

[72]

Shen W, Le S, Li Y et al. SeqKit: a cross-platform and ultrafast toolkit for FASTA/Q file manipulation. PLoS One. 2016; 11: e0163962.

[73]

Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 1999; 27: 573-80.

[74]

Reimer JJ, Turck F . Genome-wide mapping of protein-DNA interaction by chromatin immunoprecipitation and DNA microarray hybridization (ChIP-chip). Part a: ChIP-chip molecular methods. Methods Mol Biol. 2010; 631: 139-60.

[75]

Altschul SF, Madden TL, Schaffer AA et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997; 25: 3389-402.

[76]

Quinlan AR, Hall IM . BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010; 26: 841-2.

[77]

Martin M . Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnetjournal. 2011; 17: 3.

[78]

Ramirez F, Dundar F, Diehl S et al. deepTools: a flexible platform for exploring deep-sequencing data. Nucleic Acids Res. 2014; 42: W187-91.

[79]

Liu T . Use model-based analysis of ChIP-Seq (MACS) to analyze short reads generated by sequencing protein-DNA interactions in embryonic stem cells. Methods Mol Biol. 2014; 1150: 81-95.

PDF (1527KB)

112

Accesses

0

Citation

Detail

Sections
Recommended

/