A high-quality chromosome-level genome assembly reveals genetics for important traits in eggplant

Qingzhen Wei , Jinglei Wang , Wuhong Wang , Tianhua Hu , Haijiao Hu , Chonglai Bao

Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) : 153

PDF (4467KB)
Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) :153 DOI: 10.1038/s41438-020-00391-0
Article
research-article
A high-quality chromosome-level genome assembly reveals genetics for important traits in eggplant
Author information +
History +
PDF (4467KB)

Abstract

Eggplant (Solanum melongena L.) is an economically important vegetable crop in the Solanaceae family, with extensive diversity among landraces and close relatives. Here, we report a high-quality reference genome for the eggplant inbred line HQ-1315 (S. melongena-HQ) using a combination of Illumina, Nanopore and 10X genomics sequencing technologies and Hi-C technology for genome assembly. The assembled genome has a total size of ~1.17 Gb and 12 chromosomes, with a contig N50 of 5.26 Mb, consisting of 36,582 protein-coding genes. Repetitive sequences comprise 70.09% (811.14 Mb) of the eggplant genome, most of which are long terminal repeat (LTR) retrotransposons (65.80%), followed by long interspersed nuclear elements (LINEs, 1.54%) and DNA transposons (0.85%). The S. melongena-HQ eggplant genome carries a total of 563 accession-specific gene families containing 1009 genes. In total, 73 expanded gene families (892 genes) and 34 contraction gene families (114 genes) were functionally annotated. Comparative analysis of different eggplant genomes identified three types of variations, including single-nucleotide polymorphisms (SNPs), insertions/deletions (indels) and structural variants (SVs). Asymmetric SV accumulation was found in potential regulatory regions of protein-coding genes among the different eggplant genomes. Furthermore, we performed QTL-seq for eggplant fruit length using the S. melongena-HQ reference genome and detected a QTL interval of 71.29–78.26 Mb on chromosome E03. The gene Smechr0301963, which belongs to the SUN gene family, is predicted to be a key candidate gene for eggplant fruit length regulation. Moreover, we anchored a total of 210 linkage markers associated with 71 traits to the eggplant chromosomes and finally obtained 26 QTL hotspots. The eggplant HQ-1315 genome assembly can be accessed at http://eggplant-hq.cn. In conclusion, the eggplant genome presented herein provides a global view of genomic divergence at the whole-genome level and powerful tools for the identification of candidate genes for important traits in eggplant.

Cite this article

Download citation ▾
Qingzhen Wei, Jinglei Wang, Wuhong Wang, Tianhua Hu, Haijiao Hu, Chonglai Bao. A high-quality chromosome-level genome assembly reveals genetics for important traits in eggplant. Horticulture Research, 2020, 7 (1) : 153 DOI:10.1038/s41438-020-00391-0

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Chapman, M. A. The Eggplant Genome. (Springer Nature Switzerland AG, 2019).

[2]

Weese, T. L. & Bohs, L. Eggplant origins: out of Africa, into the Orient. Taxon 59, 49-56 (2010).

[3]

Daunay, M. C. et al. Genetic resources of eggplant (Solanum melongena L.) and allied species: a new challenge for molecular geneticists and eggplant breeders. Solanaceae V., Advances in Taxonomy and Utilization. (Nijmegen University Press, pp. 251-274, Nijmegen, The Netherlands, 2001).

[4]

Meyer, R. S., Karol, K. G., Little, D. P., Nee, M. H. & Litt, A. Phylogeographic relationships among Asian eggplants and new perspectives on eggplant domestication. Mol. Phylogenetics Evol. 63, 685-701 (2012).

[5]

Page, A., Gibson, J., Meyer, R. S. & Chapman, M. A. Eggplant domestication: pervasive gene flow, feralization, and transcriptomic divergence. Mol. Biol. Evol. 36, 1359-1372 (2019).

[6]

Huang, S. et al. The genome of the cucumber, Cucumis sativus L. Nat. Genet. 41, 1275-1281 (2009).

[7]

Velasco, R. et al. The genome of the domesticated apple (Malus x domestica Borkh.). Nat. Genet. 42, 833-839 (2010).

[8]

Iorizzo, M. et al. A high-quality carrot genome assembly provides new insights into carotenoid accumulation and asterid genome evolution. Nat. Genet. 48, 657-666 (2016).

[9]

The Potato Genome Sequencing Consortium . Genome sequence and analysis of the tuber crop potato. Nature 475, 189-195 (2011).

[10]

The Tomato genome Consortium . The tomato genome sequence provides insights into fleshy fruit evolution. Nature 485, 635-641 (2012).

[11]

Kim, S. et al. Genome sequence of the hot pepper provides insights into the evolution of pungency in Capsicum species. Nat. Genet. 46, 270-278 (2014).

[12]

Qin, C. et al. Whole-genome sequencing of cultivated and wild peppers provides insights into Capsicum domestication and specialization. Proc. Natl Acad. Sci. USA 111, 5135-5140 (2014).

[13]

Hirakawa, H. et al. Draft genome sequence of eggplant (Solanum melongena L.): the representative solanum species indigenous to the old world. DNA Res. 21, 649-660 (2014).

[14]

Sun, D. L. et al. Draft genome sequence of cauliflower (Brassica oleracea L. var. botrytis) provides new insights into the C genome in Brassica species. Hortic. Res. 6, 82 (2019).

[15]

Li, M. Y. et al. The genome sequence of celery (Apium graveolens L.), an important leaf vegetable crop rich in apigenin in the Apiaceae family. Hortic. Res. https://doi.org/10.1038/s41438-019-0235-2 (2020).

[16]

Barchi, L. et al. A chromosome-anchored eggplant genome sequence reveals key events in Solanaceae evolution. Sci. Rep. 9, 11769 (2019).

[17]

Song, B. et al. Draft genome sequence of Solanum aethiopicum provides insights into disease resistance, drought tolerance, and the evolution of the genome. GigaScience 8, 1-16 (2019).

[18]

The Arabidopsis Genome Initiative . Analysis of the genome sequence of the flowering plant Arabidopsis thaliana. Nature 408, 796-815 (2000).

[19]

Yu, J. et al. A draft sequence of the rice genome (Oryza sativa L. ssp. indica). Science 296, 79-92 (2002).

[20]

Xia, E. H. et al. The tea tree genome provides insights into tea flavor and independent evolution of caffeine biosynthesis. Mol. Plant 10, 866-877 (2017).

[21]

Cécile, C. et al. Source of resistance against Ralstonia solanacearum in fertile somatic hybrids of eggplant (Solanum melongena L.) with Solanum aethiopicum L. Plant Sci. 160, 301-313 (2001).

[22]

Gisbert, C. et al. Eggplant relatives as sources of variation for developing new rootstocks: effects of grafting on eggplant yield and fruit apparent quality and composition. Sci. Horti. 128, 14-22 (2011).

[23]

Barchi, L. et al. A RAD tag derived marker based eggplant linkage map and the location of QTLs determining anthocyanin pigmentation. PLoS ONE 7, e43740 (2012).

[24]

Barchi, L. et al. QTL analysis reveals new eggplant loci involved in resistance to fungal wilts. Euphytica 214, 20 (2018).

[25]

Cericola, F. et al. Linkage disequilibrium and genome-wide association analysis for anthocyanin pigmentation and fruit color in eggplant. BMC Genomics 15, 896 (2014).

[26]

Portis, E. et al. QTL mapping in eggplant reveals clusters of yield-related loci and orthology with the tomato genome. PLoS ONE 9, e89499 (2014).

[27]

Portis, E. et al. Association mapping for fruit, plant and leaf morphology traits in eggplant. PLoS ONE 10, e0135200 (2015).

[28]

Toppino, L. et al. Mapping quantitative trait loci affecting biochemical and morphological fruit properties in eggplant (Solanum melongena L.). Front. Plant Sci. 4, 256 (2016).

[29]

Miyatake, K. et al. Detailed mapping of a resistance locus against Fusarium wilt in cultivated eggplant (Solanum melongena L.). Theor. Appl. Genet. 129, 357-367 (2016).

[30]

Wei, Q. Z. et al. Construction of a SNP-Based genetic map using SLAF-Seq and QTL analysis of morphological traits in eggplant. Front. Genet. 11, 178 (2020).

[31]

Bolger, M. E., Arsova, B. & Usadel, B. Plant genome and transcriptome annotations: from misconceptions to simple solutions. Brief. Bioinform. 19, 437-449 (2018).

[32]

Jiao, W. B. & Schneeberger, K. The impact of third generation genomic technologies on plant genome assembly. Curr. Opin. Plant Biol. 36, 64-70 (2017).

[33]

Berlin, K. et al. Assembling large genomes with single molecule sequencing and locality-sensitive hashing. Nat. Biotechnol. 33, 623-630 (2015).

[34]

Hirsch, C. N. et al. Draft assembly of elite inbred line PH207 provides insights into genomic and transcriptome diversity in maize. Plant Cell 28, 2700-2714 (2016).

[35]

Maximilian, H. W. et al. De novo assembly of a new Solanum pennellii accession using nanopore sequencing. Plant Cell 29, 2336-2348 (2017).

[36]

Reyes-Chin-Wo, S. et al. Genome assembly with in vitro proximity ligation data and whole-genome triplication in lettuce. Nat. Commun. 8, 14953 (2017).

[37]

Sierro, N. et al. The tobacco genome sequence and its comparison with those of tomato and potato. Nat. Commun. 5, 3833 (2014).

[38]

Bombarely, A. et al. Insight into the evolution of the Solanaceae from the parental genomes of Petunia hybrid. Nat. Plants 2, 16074 (2016).

[39]

Carvalho, C. M. B. & Lupski, J. R. Mechanisms underlying structural variant formation in genomic disorders. Nat. Rev. Genet. 17, 224-238 (2016).

[40]

Cook, D. E. et al. Missing heritability and strategies for finding the underlying causes of complex disease. Nat. Rev. Genet. 11, 446-450 (2010).

[41]

Chakraborty, M. et al. Hidden genetic variation shapes the structure of functional elements in Drosophila. Nat. Genet. 50, 20-25 (2018).

[42]

Yin, D. M. et al. Comparison of Arachis monticola with diploid and cultivated tetraploid genomes reveals asymmetric subgenome evolution and improvement of peanut. Adv. Sci. 7, 1901672 (2019).

[43]

Prohens, J. et al. Characterization of interspecific hybrids and first backcross generations from crosses between two cultivated eggplants (Solanum melongena and S. aethiopicum Kumba group) and implications for eggplant breeding. Euphytica 186, 517-538 (2012).

[44]

Collonnier, C. et al. Source of resistance against Ralstonia solanacearum in fertile somatic hybrids of eggplant (Solanum melongena L.) with Solanum aethiopicum L. Plant Sci. 160, 301-313 (2001).

[45]

Frary, A. et al. QTL hotspots in eggplant (Solanum melongena) detected with a high resolution map and CIM analysis. Euphytica 197, 211-228 (2014).

[46]

Fukuoka, H. et al. Development of gene-based markers and construction of an integrated linkage map in eggplant by using Solanum orthologous (SOL) gene sets. Theor. Appl. Genet. 125, 47-56 (2012).

[47]

Ruan, J. & Li, H. Fast and accurate long-read assembly with wtdbg2. Nat. Methods 17, 1-4 (2020).

[48]

Vaser, R., Sovic, I., Nagarajan, N. & Sikic, M . Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 27, 737-746 (2017).

[49]

Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760 (2009).

[50]

Mostovoy, Y. et al. A hybrid approach for de novo human genome sequence assembly and phasing. Nat. Methods 13, 587-590 (2016).

[51]

Li, H. et al. 1000 Genome Project Data Processing Subgroup. The Sequence Alignment/Map format and SAMtools. Bioinformatics 25, 2078-2079 (2009).

[52]

Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119-1125 (2013).

[53]

Simão, F. A. et al. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210-3212 (2015).

[54]

Grabherr, M. G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat. Biotechnol. 29, 644-652 (2011).

[55]

Kim, D. et al. TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions. Genome Biol. 14, R36 (2013).

[56]

Trapnell, C. et al. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat. Biotechnol. 28, 511-515 (2010).

[57]

Haas, B. J. et al. Automated eukaryotic gene structure annotation using evidencemodeler and the program to assemble spliced alignments. Genome Biol. 9, R7 (2008).

[58]

Campbell, M. A., Hass, B. J., Hamilton, J. P., Mount, S. M. & Buell, C. R. Comprehensive analysis of alternative splicing in rice and comparative analyses with Arabidopsis. BMC Genom. 7, 1-7 (2006).

[59]

Stanke, M., Steinkamp, R., Waack, S. & Morgenstern, B. AUGUSTUS: a web server for gene finding in eukaryotes. Nucleic Acids Res. 32, 309-312 (2004).

[60]

Parra, G., Blanco, E. & Guigó, R. GeneID in Drosophila. Genome Res. 10, 511 (2000).

[61]

Aggarwal, G. & Ramaswamy, R. Ab initio gene identification: prokaryote genome annotation with GeneScan and GLIMMER. J. Biosci. 27, 7-14 (2002).

[62]

Majoros, W. H., Pertea, M. & Salzberg, S. L. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics 20, 2878-2879 (2004).

[63]

Bromberg, Y. & Rost, B. SNAP: predict effect of non-synonymous polymorphisms on function. Nucleic Acids Res. 35, 3823 (2007).

[64]

Smit, A. F. A., Hubley, R. & Green, P. RepeatMasker Open-4.0. http://www.repeatmasker.org (1996-2013).

[65]

Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, 265-268 (2007).

[66]

Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat families in large genomes. Bioinformatics 21, 351-358 (2005).

[67]

Smit, A. F. A. & Hubley, R. RepeatModeler Open-1.0. http://www.repeatmasker.org (2008-2015).

[68]

Altschul, S. F. et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 25, 3389-3402 (1997).

[69]

Birney, E., Clamp, M. & Durbin, R. GeneWise and Genomewise. Genome Res. 14, 988-995 (2004).

[70]

Gish, W. & States, D. J. Identification of protein coding regions by database similarity search. Nat. Genet. 3, 266-272 (1993).

[71]

Gouzy, J. et al. XDOM, a graphical tool to analyse domain arrangements in any set of protein sequences. CABIOS 13, 601-608 (1997).

[72]

Finn, R. D. et al. Pfam: the protein families database. Nat. Genet. 42, 222-230 (2014).

[73]

Letunic, I. et al. SMART 4.0: towards genomic data integration. Nucleic Acids Res. 32, 142-144 (2004).

[74]

Mi, H. Y., Muruganujan, A. & Thomas, P. D. PANTHER in 2013: modeling the evolution of gene function, and other gene attributes, in the context of phylogenetic trees. Nucleic Acids Res. 41, 377-386 (2012).

[75]

Sigrist, C. J. A. et al. New and continuing developments at PROSITE. Nucleic Acids Res. 41, 344-347 (2013).

[76]

Hunter, S. et al. InterPro: the integrative protein signature database. Nucleic Acids Res. 37, 211-215 (2009).

[77]

Ashburner, M. et al. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nat. Genet. 25, 25-29 (2000).

[78]

Kanehisa, M. et al. Data, information, knowledge and principle: back to metabolism in KEGG. Nucleic Acids Res. 42, 199-205 (2014).

[79]

Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res. 25, 955-964 (1997).

[80]

Griffiths-Jones, S. et al. Rfam: annotating non-codin RNAs in complete genomes. Nucleic Acids Res. 33, 121-124 (2005).

[81]

Nawrocki, E. P., Kolbe, D. L. & Eddy, S. R. Infernal 1.0: inference of RNA alignments. Bioinformatics 25, 1335 (2009).

[82]

Han, M. V., Thomas, G. W., Lugo-Martinez, J. & Hahn, M. W. Estimating gene gain and loss rates in the presence of error in genome assembly and annotation using CAFE 3. Mol. Biol. Evol. 30, 1987-1997 (2013).

[83]

Edgar, R. C. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 32, 1792-1797 (2004).

[84]

Stamatakis, A. RAxML version 8: a tool for phylogenetic analysis and postanalysis of large phylogenies. Bioinformatics 30, 1312-1313 (2014).

[85]

Yang, Z. H. PAML 4: phylogenetic analysis by maximum likelihood. Mol. Biol. Evol. 24, 1586-1591 (2007).

[86]

McKenna, A. et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 20, 1297-1303 (2010).

[87]

Chen, K. et al. Breakdancer: an algorithm for high-resolution mapping of genomic structural variation. Nat. Methods 6, 677-681 (2009).

[88]

Takagi, H. et al. QTL-seq: rapid mapping of quantitative trait loci in rice by whole genome resequencing of DNA from two bulked populations. Plant J. 74, 174-183 (2013).

[89]

Wang, K., Li, M. & Hakonarson, H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164 (2010).

PDF (4467KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/