The genome of Chinese flowering cherry (Cerasus serrulata) provides new insights into Cerasus species

Xian-Gui Yi , Xia-Qing Yu , Jie Chen , Min Zhang , Shao-Wei Liu , Hong Zhu , Meng Li , Yi-Fan Duan , Lin Chen , Lei Wu , Shun Zhu , Zhong-Shuai Sun , Xin-Hong Liu , Xian-Rong Wang

Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) : 165

PDF (2461KB)
Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) :165 DOI: 10.1038/s41438-020-00382-1
Article
research-article
The genome of Chinese flowering cherry (Cerasus serrulata) provides new insights into Cerasus species
Author information +
History +
PDF (2461KB)

Abstract

Cerasus serrulata is a flowering cherry germplasm resource for ornamental purposes. In this work, we present a de novo chromosome-scale genome assembly of C. serrulata by the use of Nanopore and Hi-C sequencing technologies. The assembled C. serrulata genome is 265.40 Mb across 304 contigs and 67 scaffolds, with a contig N50 of 1.56 Mb and a scaffold N50 of 31.12 Mb. It contains 29,094 coding genes, 27,611 (94.90%) of which are annotated in at least one functional database. Synteny analysis indicated that C. serrulata and C. avium have 333 syntenic blocks composed of 14,072 genes. Blocks on chromosome 01 of C. serrulata are distributed on all chromosomes of C. avium, implying that chromosome 01 is the most ancient or active of the chromosomes. The comparative genomic analysis confirmed that C. serrulata has 740 expanded gene families, 1031 contracted gene families, and 228 rapidly evolving gene families. By the use of 656 single-copy orthologs, a phylogenetic tree composed of 10 species was constructed. The present C. serrulata species diverged from Prunus yedoensis ~17.34 million years ago (Mya), while the divergence of C. serrulata and C. avium was estimated to have occurred ∼21.44 Mya. In addition, a total of 148 MADS-box family gene members were identified in C. serrulata, accompanying the loss of the AGL32 subfamily and the expansion of the SVP subfamily. The MYB and WRKY gene families comprising 372 and 66 genes could be divided into seven and eight subfamilies in C. serrulata, respectively, based on clustering analysis. Nine hundred forty-one plant disease-resistance genes (R-genes) were detected by searching C. serrulata within the PRGdb. This research provides high-quality genomic information about C. serrulata as well as insights into the evolutionary history of Cerasus species.

Cite this article

Download citation ▾
Xian-Gui Yi, Xia-Qing Yu, Jie Chen, Min Zhang, Shao-Wei Liu, Hong Zhu, Meng Li, Yi-Fan Duan, Lin Chen, Lei Wu, Shun Zhu, Zhong-Shuai Sun, Xin-Hong Liu, Xian-Rong Wang. The genome of Chinese flowering cherry (Cerasus serrulata) provides new insights into Cerasus species. Horticulture Research, 2020, 7 (1) : 165 DOI:10.1038/s41438-020-00382-1

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Du, H. K. Practical methods for rapid seed germination from seed coat-imposed dormancy of Prunus yedoensis. Scientia Hortic. 243, 451-456 (2019).

[2]

Balsamo, R. A. et al. Leaf biomechanics, morphology, and anatomy of the deciduous mesophyte Prunus serrulata (Rosaceae) and the evergreen sclerophyllous shrub Heteromelesarbutifolia (Rosaceae). Am. J. Bot. 90, 72-77 (2003).

[3]

Liu, Z. X. et al. Development of stamens and carpels in single and double flowers of Cerasus serrulata. J. Beijing For. Univ. 32, 486-491 (2010).

[4]

Li, C. L. et al. Cerasus in Flora of China. Science Press. 9, 404-420 (2003).

[5]

Yi, X. G. The variation and phylogeography of Cerasus serrulata Mill. populations. J. Nanjing For. Univ. 14, 166-172 (2018).

[6]

Wang, X. R. An illustrated monograph of cherry cultivars in China. Science Press. 12, 24-28 (2014).

[7]

Ma, H., Olsen, R., Pooler, M. & Kramer, M. Evaluation of flowering cherry species, hybrids, and cultivars using simple sequence repeat markers. SocHort Science. 134, 435-444 (2009).

[8]

Knight, R. Abstract bibliography of fruit breeding and genetics to 1965, Prunus. Commonwealth Agricultural Bureau. 3, 752-824 (1969).

[9]

Iwatsuki, K., Boufford, D. E. & Ohba, H. Flora of Japan. Science Press 2, 435-148 (2001).

[10]

Meng, L. I. et al. Numeric and structural characteristics of Cerasus serrulata population around the high-elevation wetlands of Dayangshan. J. Nanjing For. Univ. 37, 40-44 (2013).

[11]

Hong, Z. et al. Application of the molecular marker technology to Cerasus Mill. (Rosaceae). World For. Res. 31, 16-24 (2018).

[12]

Shirasawa, K. et al. The genome sequence of sweet cherry (Prunus avium) for use in genomics-assisted breeding. DNA Research. 24, 499-508 (2017).

[13]

Verde, I. et al. The high-quality draft genome of peach (Prunus persica) identifies unique patterns of genetic diversity, domestication and genome evolution. Nat. Genet. 45, 487-494 (2013).

[14]

Velasco, R. et al. The genome of the domesticated apple (Malus x domestica). Nat. Genet. 42, 833-839 (2010).

[15]

Velasco, D. et al. Evolutionary genomics of peach and almond domestication. G3 32, 116-121 (2016).

[16]

Shulaev, V. et al. The genome of woodland strawberry (Fragaria vesca). Nat. Genet. 43, 109-116 (2011).

[17]

Jiang, F. et al. The apricot (Prunus armeniaca L.) genome elucidates Rosaceae evolution and beta-carotenoid synthesis. Hortic. Res. 6, 128-134 (2019).

[18]

Vanburen, R. et al. The genome of black raspberry (Rubus occidentalis). Plant J. 87, 535-547 (2016).

[19]

Seunghoon, B. et al. Draft genome sequence of wild Prunus yedoensis reveals massive inter-specific hybridization between sympatric flowering cherries. Genome Biol. 19, 45-51 (2018).

[20]

Zhang, Q. et al. The genome of Prunus mume. Nat. Commun. 3, 1318-1325 (2012).

[21]

Lin, W. et al. Characterization of the complete chloroplast genome of Chinese rose, Rosa chinensis (Rosaceae: Rosa). Mitochondrial DNA B. 4, 51-58 (2019).

[22]

Hibrand, S. L. et al. A high-quality genome sequence of Rosa chinensis to elucidate ornamental traits. Nat. Plants 4, 36-43 (2018).

[23]

Fei, C. et al. Evolutionary analysis of MIKCc-type MADS-box genes in gymnosperms and angiosperms. Front. Plant Sci. 8, 895-899 (2017).

[24]

Wells, C. E. et al. A genome-wide analysis of MADS-box genes in peach. BMC Plant Biol. 15, 41-46 (2015).

[25]

Tian, Y. et al. Genome-wide identification and analysis of the MADS-box gene family in apple. Gene 555, 277-290 (2014).

[26]

Katoh, K., Rozewicki, J. & Yamada, K. D. MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization. Brief. Bioinform. 4, 20-26 (2019).

[27]

Morgan, N. P., Paramvir, S. D. & Adam, P. A. FastTree2-approximately maximum-likelihood trees for large alignments. PLOS ONE 3, 5-12 (2010).

[28]

Bayliss, S. C. et al. The use of Oxford Nanopore native barcoding for complete genome assembly. GigaScience 6, 1-6 (2017).

[29]

Goodwin, S. et al. Oxford Nanopore sequencing, hybrid error correction, and de novo assembly of a eukaryotic genome. Genome Res. 25, 1750-1758 (2015).

[30]

Vaser, R. et al. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 5, 27-32 (2017).

[31]

Brummitt, R. K. et al. The species plantarum project, an international collaborative initiative for higher plant taxonomy. Taxon 50, 1217-1230 (2001).

[32]

Miller, P. containing the methods of cultivating and improving the kitchen, fruit and flower garden, as also the physick garden, wilderness, conservatory, and vineyard. Gardeners Dictionary 3, 576-614 (1753).

[33]

Innan, H., Terauchi, R., Miyashita, N. T. & Tsunewaki, K. DNA fingerprinting study on the intraspecific variation and the origin of Prunus yedoensis (Someiyoshino). Jpn. J. Genet. 70, 185-196 (1995).

[34]

Kato, S. et al. Origins of Japanese flowering cherry (Prunus subgenus Cerasus) cultivars revealed using nuclear SSR markers. Tree Genet. Genomes 10, 477-487 (2014).

[35]

Chen, S. et al. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, 884-890 (2018).

[36]

Altschul, S. F. et al. Basic local alignment search tool. J. Mol. Biol. 215, 403-410 (1990).

[37]

Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 27, 764-770 (2011).

[38]

Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722-736 (2017).

[39]

Lin, Y. et al. Assembly of long error-prone reads using de Bruijn graphs. Proc. Natl Acad. Sci. USA 113, 52-56 (2016).

[40]

Jue, R. smartdenovo: ultra-fast de novo assembler using long noisy reads. https://github.com/ruanjue/smartdenovo (2015).

[41]

Chakraborty, M. et al. Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage. Nucleic Acids Res. 44, 147-151 (2016).

[42]

Kurtz, S. et al. Versatile and open software for comparing large genomes. Genome Biol. 5, 12-15 (2004).

[43]

Walker, B. J. et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS ONE 9, 29-32 (2014).

[44]

Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. Genomics. 3, 76-78 (2013).

[45]

Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16, 1-11 (2015).

[46]

Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119-1125 (2013).

[47]

Xu, G. C. et al. LR Gapcloser: a tiling path-based gap closer that uses long reads to complete genome assembly. Gigascience 8, 157-160 (2018).

[48]

Simão, F. A. et al. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 19-23 (2015).

[49]

Parra, G., Bradnam, K. & Korf, I. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics 23, 1061-1067 (2007).

[50]

Li, H. et al. The sequence alignment/map format and SAMtools. Bioinformatics 25, 2078-2079 (2009).

[51]

Xu, Z. & Wang, H. LTR FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, 265-268 (2007).

[52]

Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat families in large genomes. Bioinformatics 21, 351-358 (2005).

[53]

Edgar, R. C. & Myers, E. W. PILER: identification and classification of genomic repeats. Bioinformatics 21, 152-158 (2005).

[54]

Wicker, T. et al. A unified classification system for eukaryotic transposable elements. Nat. Rev. Genet. 8, 973-982 (2007).

[55]

Smit, A., Hubley. R. & Green, P. RepeatMasker Open-4.0 (2013-2015). http://repeatmasker.org (2017).

[56]

Burge, C. & Karlin, S. Prediction of complete gene structures in human genomic DNA. J. Mol. Biol. 268, 78-94 (1997).

[57]

Stanke, M. & Waack, S. Gene prediction with a hidden Markov model and a new intron submodel. Bioinformatics 19, 215-225 (2003).

[58]

Majoros, W. H. et al. Two open source ab initio eukaryotic gene-finders. Bioinformatics 20, 2878-2879 (2004).

[59]

Blanco, E., Parra, G. & Guigó, R. Using geneid to identify genes. Curr. Protoc Bioinform. 18, 1-28 (2007).

[60]

Korf, I. Gene finding in novel genomes. BMC Bioinform. 5, 59-62 (2004).

[61]

Meyerowitz, E. M. Arabidopsis thaliana. Ann. Rev. Genet. 21, 93-111 (2003).

[62]

Keilwagen, J. et al. Using intron position conservation for homology-based gene prediction. Nucleic Acids Res. 44, 89-95 (2016).

[63]

Pertea, M. et al. Transcript-level expression analysis of RNA-seq experiments with HISAT, StringTie and Ballgown. Nat. Protoc. 11, 1650-1658 (2016).

[64]

Haas, B. J. & Papanicolaou, A. TransDecoder (Find Coding Regions Within Transcripts). http://transdecoder.github.io (2015).

[65]

Tang, S., Lomsadze, A. & Borodovsky, M. Identification of protein coding regions in RNA transcripts. Nucleic Acids Res. 43, 78-85 (2015).

[66]

Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 9, 7-15 (2008).

[67]

Campbell, M. A. et al. Comprehensive analysis of alternative splicing in rice and comparative analyses with Arabidopsis. BMC Genomics 7, 324-327 (2006).

[68]

She, R. et al. genBlastG: using BLAST searches to build homologous gene models. Bioinformatics 27, 2141-2143 (2011).

[69]

Nawrocki, E. P. & Eddy, S. R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29, 2933-2935 (2013).

[70]

Griffiths, J. S. et al. miRBase: microRNA sequences, targets and gene nomenclature. Nucleic Acids Res. 34, 140-144 (2006).

[71]

Griffiths, J. S. et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 33, 121-124 (2005).

[72]

Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNAgenes in genomic sequence. Nucleic Acids Res. 25, 955-964 (1997).

[73]

Tatusov, R. L. et al. The COG database: new developments in phylogenetic classification of proteins from complete genomes. Nucleic Acids Res. 29, 22-28 (2001).

[74]

Kanehisa, M. & Goto, S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 28, 27-30 (2000).

[75]

Boeckmann, B. et al. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Res. 31, 365-370 (2003).

[76]

Marchler, B. A. et al. CDD: a conserved domain database for the functional annotation of proteins. Nucleic Acids Res. 39, 225-229 (2011).

[77]

El, G. S. et al. The Pfam protein families database in 2019. Nucleic Acids Res. 47, 427-432 (2019).

[78]

Dimmer, E. C. et al. The UniProt-GO annotation database in 2011. Nucleic Acids Res. 40, 565-570 (2012).

[79]

Eddy, S. R., Mitchison, G. & Durbin, R. Maximum discrimination hidden Markov models of sequence consensus. J. Comput. Biol. 2, 9-23 (1995).

[80]

Conesa, A. et al. Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research. Bioinformatics 21, 3674-3676 (2005).

[81]

De, B. T. et al. CAFE: a computational tool for the study of gene family evolution. Bioinformatics 22, 1269-1271 (2006).

[82]

Li, L., Stoeckert, C. J. & Roos, D. S. OrthoMCL: identification of ortholog groups for eukaryotic genomes. Genome Res. 13, 2178-2189 (2003).

[83]

Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol. Biol. Evol. 30, 772-780 (2013).

[84]

Capella, G. S., Silla, M. & Gabaldón, T. trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics 25, 1972-1973 (2009).

[85]

Kalyaanamoorthy, S., Minh, B. Q., Wong, T. K., Von, H. A. & Jermiin, L. S. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat. Methods 14, 587-592 (2017).

[86]

Edgar, R. C. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 32, 1792-1797 (2004).

[87]

Guindon, S. et al. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Syst. Biol. 59, 307-321 (2010).

[88]

Yang, Z. H. PAML 4: phylogenetic analysis by maximum likelihood. Mol. Biol. Evol. 24, 1586-1591 (2007).

PDF (2461KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/