The high-quality sequencing of the Brassica rapa ‘XiangQingCai’ genome and exploration of genome evolution and genes related to volatile aroma

Zhaokun Liu , Yanhong Fu , Huan Wang , Yanping Zhang , Jianjun Han , Yingying Wang , Shaoqin Shen , Chunjin Li , Mingmin Jiang , Xuemei Yang , Xiaoming Song

Horticulture Research ›› 2023, Vol. 10 ›› Issue (10) : 187

PDF (6667KB)
Horticulture Research ›› 2023, Vol. 10 ›› Issue (10) :187 DOI: 10.1093/hr/uhad187
Article
research-article
The high-quality sequencing of the Brassica rapa ‘XiangQingCai’ genome and exploration of genome evolution and genes related to volatile aroma
Author information +
History +
PDF (6667KB)

Abstract

‘Vanilla’ (XQC, brassica variety chinensis) is an important vegetable crop in the Brassica family, named for its strong volatile fragrance. In this study, we report the high-quality chromosome-level genome sequence of XQC. The assembled genome length was determined as 466.11 Mb, with an N50 scaffold of 46.20 Mb. A total of 59.50% repetitive sequences were detected in the XQC genome, including 47 570 genes. Among all examined Brassicaceae species, XQC had the closest relationship with B. rapa QGC (‘QingGengCai’) and B. rapa Pakchoi. Two whole-genome duplication (WGD) events and one recent whole-genome triplication (WGT) event occurred in the XQC genome in addition to an ancient WGT event. The recent WGT was observed to occur during 21.59–24.40 Mya (after evolution rate corrections). Our findings indicate that XQC experienced gene losses and chromosome rearrangements during the genome evolution of XQC. The results of the integrated genomic and transcriptomic analyses revealed critical genes involved in the terpenoid biosynthesis pathway and terpene synthase (TPS) family genes. In summary, we determined a chromosome-level genome of B. rapa XQC and identified the key candidate genes involved in volatile fragrance synthesis. This work can act as a basis for the comparative and functional genomic analysis and molecular breeding of B. rapa in the future.

Cite this article

Download citation ▾
Zhaokun Liu, Yanhong Fu, Huan Wang, Yanping Zhang, Jianjun Han, Yingying Wang, Shaoqin Shen, Chunjin Li, Mingmin Jiang, Xuemei Yang, Xiaoming Song. The high-quality sequencing of the Brassica rapa ‘XiangQingCai’ genome and exploration of genome evolution and genes related to volatile aroma. Horticulture Research, 2023, 10 (10) : 187 DOI:10.1093/hr/uhad187

登录浏览全文

4963

注册一个新账户 忘记密码

Acknowledgements

This work was supported by the Suzhou Agricultural Science and Technology Innovation project (SNG2020065; SNG2020045), Suzhou Municipal Bureau of Agriculture and Rural Affairs, the National Natural Science Foundation of China (32172583), and the Natural Science Foundation of Hebei (C2021209005). The genome sequencing was performed in the Novogene Corporation.

Author contributions

Z.L. was responsible for the project initiation. Z.L. and X.S. supervised and managed the project and research. Experiments and analyses were designed by Z.L., X.S., H.W., and Y.Z. Bioinformatic analyses were led by X.S., Z.L., Y.F., Y.Z., and S.S. The manuscript was written and revised by Z.L., X. S., Y.F., H.W., and Y.Z. All authors read and revised the manuscript.

Data availability

The XQC genome sequence and RNA-seq datasets have been deposited in the Genome Sequence Archive [85] of the BIG Data Center [86], under accession numbers CRA010486 and CRA010488. They are publicly accessible at http://bigd.big.ac.cn/gsa. The genome sequences and annotation of XQC can be downloaded from the TBGR database (http://www.tbgr.org.cn) with the Genome ID of Pakchoi-XQC-v1.0 [30].

Conflict of interest statement

The authors declare no competing interests.

References

[1]

Song X, Wei Y, Xiao D et al. Brassica carinata genome characterization clarifies U’s triangle model of evolution and polyploidy in brassica. Plant Physiol. 2021; 186: 388-406

[2]

Nagaharu U . Genome analysis in brassica with special reference to the experimental formation of B. napus and peculiar mode of fertilication . Jpn J Bot. 1935; 7: 389-452

[3]

Wang X, Wang H, Wang J et al. The genome of the mesopolyploid crop species Brassica rapa. Nat Genet. 2011; 43: 1035-9

[4]

Cai C, Wang X, Liu B et al. Brassica rapa genome 2.0: a reference upgrade through sequence reassembly and gene reannotation. Mol Plant. 2017; 10: 649-51

[5]

Zhang L, Cai X, Wu J et al. Improved Brassica rapa reference genome by single-molecule sequencing and chromosome conformation capture technologies . Hortic Res. 2018; 5: 50

[6]

Zhang Z, Guo J, Cai X et al. Improved reference genome annotation of Brassica rapa by Pacific biosciences RNA sequencing . Front Plant Sci. 2022; 13: 841618

[7]

Yang Z, Jiang Y, Gong J et al. R gene triplication confers European fodder turnip with improved clubroot resistance. Plant Biotechnol J. 2022; 20: 1502-17

[8]

Li Y, Liu GF, Ma LM et al. A chromosome-level reference genome of non-heading Chinese cabbage [ Brassica campestris (syn. Brassica rapa) ssp. chinensis] . Hortic Res. 2020; 7: 212

[9]

Li P, Su T, Zhao X et al. Assembly of the non-heading pak choi genome and comparison with the genomes of heading Chinese cabbage and the oilseed yellow sarson. Plant Biotechnol J. 2021; 19: 966-76

[10]

Xu H, Wang C, Shao G et al. The reference genome and full-length transcriptome of pakchoi provide insights into cuticle formation and heat adaption. Hortic Res. 2022; 9: uhac123

[11]

Zhang L, Liang J, Chen H et al. A near-complete genome assembly of Brassica rapa provides new insights into the evolution of centromeres . Plant Biotechnol J. 2023; 21: 1022-32

[12]

Liu S, Liu Y, Yang X et al. The Brassica oleracea genome reveals the asymmetrical evolution of polyploid genomes . Nat Commun. 2014; 5: 3930

[13]

Parkin IA, Koh C, Tang H et al. Transcriptome and methylome profiling reveals relics of genome dominance in the mesopolyploid Brassica oleracea. Genome Biol. 2014; 15: R77

[14]

Sun D, Wang C, Zhang X et al. Draft genome sequence of cauliflower (Brassica oleracea L. var. botrytis) provides new insights into the C genome in Brassica species . Hortic Res. 2019; 6: 82

[15]

Lv H, Wang Y, Han F et al. A high-quality reference genome for cabbage obtained with SMRT reveals novel genomic features and evolutionary characteristics. Sci Rep. 2020; 10: 12394

[16]

Guo N, Wang S, Gao L et al. Genome sequencing sheds light on the contribution of structural variants to Brassica oleracea diversification . BMC Biol. 2021; 19: 93

[17]

Cai X, Wu J, Liang J et al. Improved Brassica oleracea JZS assembly reveals significant changing of LTR-RT dynamics in different morphotypes . Theor Appl Genet. 2020; 133: 3187-99

[18]

Perumal S, Koh CS, Jin L et al. A high-contiguity Brassica nigra genome localizes active centromeres and defines the ancestral Brassica genome . Nat Plants. 2020; 6: 929-41

[19]

Chalhoub B, Denoeud F, Liu S et al. Plant genetics. Early allopolyploid evolution in the post-Neolithic Brassica napus oilseed genome . Science. 2014; 345: 950-3

[20]

Bayer PE, Hurgobin B, Golicz AA et al. Assembly and comparison of two closely related Brassica napus genomes . Plant Biotechnol J. 2017; 15: 1602-10

[21]

Sun F, Fan G, Hu Q et al. The high-quality genome of Brassica napus cultivar ’ZS11’ reveals the introgression history in semi-winter morphotype . Plant J. 2017; 92: 452-68

[22]

Zou J, Mao L, Qiu J et al. Genome-wide selection footprints and deleterious variations in young Asian allotetraploid rapeseed. Plant Biotechnol J. 2019; 17: 1998-2010

[23]

Song JM, Guan Z, Hu J et al. Eight high-quality genomes reveal pan-genome architecture and ecotype differentiation of Brassica napus. Nat Plants. 2020; 6: 34-45

[24]

Rousseau-Gueutin M, Belser C, da Silva C et al. Long-read assembly of the Brassica napus reference genome Darmor-bzh . Gigascience. 2020; 9: giaa137

[25]

Chen X, Tong C, Zhang X et al. A high-quality Brassica napus genome reveals expansion of transposable elements, subgenome evolution and disease resistance . Plant Biotechnol J. 2021; 19: 615-30

[26]

Lee H, Chawla HS, Obermeier C et al. Chromosome-scale assembly of winter oilseed rape Brassica napus. Front Plant Sci. 2020; 11: 496

[27]

Yim WC, Swain ML, Ma D et al. The final piece of the triangle of U: evolution of the tetraploid Brassica carinata genome . Plant Cell. 2022; 34: 4143-72

[28]

Yang J, Liu D, Wang X et al. The genome sequence of allopolyploid Brassica juncea and analysis of differential homoeolog gene expression influencing selection . Nat Genet. 2016; 48: 1225-32

[29]

Paritosh K, Yadava SK, Singh P et al. A chromosome-scale assembly of allotetraploid Brassica juncea (AABB) elucidates comparative architecture of the a and B genomes . Plant Biotechnol J. 2021; 19: 602-14

[30]

Liu Z, Li N, Yu T et al. The Brassicaceae genome resource (TBGR): a comprehensive genome platform for Brassicaceae plants. Plant Physiol. 2022; 190: 226-37

[31]

Yu T, Ma X, Liu Z et al. TVIR: a comprehensive vegetable information resource database for comparative and functional genomic studies. Hortic Res. 2022; 9: uhac213

[32]

Wu J, Liang J, Lin R et al. Investigation of brassica and its relative genomes in the post-genomics era. Hortic Res. 2022; 9: uhac182

[33]

Cai X, Chang L, Zhang T et al. Impacts of allopolyploidization and structural variation on intraspecific diversification in Brassica rapa. Genome Biol. 2021; 22: 166

[34]

Aubourg S, Lecharny A, Bohlmann J . Genomic analysis of the terpenoid synthase (AtTPS) gene family of Arabidopsis thaliana. Mol Genet Genomics. 2002; 267: 730-45

[35]

Jaillon O, Aury JM, Noel B et al. The grapevine genome sequence suggests ancestral hexaploidization in major angiosperm phyla. Nature. 2007; 449: 463-7

[36]

Belser C, Istace B, Denis E et al. Chromosome-scale assemblies of plant genomes using nanopore long reads and optical maps. Nat Plants. 2018; 4: 879-87

[37]

Song X, Wang J, Li N et al. Deciphering the high-quality genome sequence of coriander that causes controversial feelings. Plant Biotechnol J. 2020; 18: 1444-56

[38]

Song X, Liu H, Shen S et al. Chromosome-level Pepino genome provides insights into genome evolution and anthocyanin biosynthesis in Solanaceae. Plant J. 2022; 110: 1128-43

[39]

Song X, Sun P, Yuan J et al. The celery genome sequence reveals sequential paleo-polyploidizations, karyotype evolution and resistance gene reduction in apiales. Plant Biotechnol J. 2021; 19: 731-44

[40]

Marcais G, Kingsford C . A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 2011; 27: 764-70

[41]

Cheng H, Concepcion GT, Feng X et al. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021; 18: 170-5

[42]

Wingett S, Ewels P, Furlan-Magaril M et al. HiCUP: pipeline for mapping and processing hi-C data. F1000Res. 2015; 4: 1310

[43]

Shen S, Li N, Wang Y et al. High-quality ice plant reference genome analysis provides insights into genome evolution and allows exploration of genes involved in the transition from C3 to CAM pathways. Plant Biotechnol J. 2022; 20: 2107-22

[44]

Zhang X, Zhang S, Zhao Q et al. Assembly of allele-aware, chromosomal-scale autopolyploid genomes based on hi-C data. Nat Plants. 2019; 5: 833-45

[45]

Durand NC, Shamim MS, Machol I et al. Juicer provides a one-click system for analyzing loop-resolution hi-C experiments. Cell Syst. 2016; 3: 95-8

[46]

Parra G, Bradnam K, Korf I . CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics. 2007; 23: 1061-7

[47]

Manni M, Berkeley MR, Seppey M et al. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 2021; 38: 4647-54

[48]

Li H, Durbin R . Fast and accurate short read alignment with burrows-wheeler transform. Bioinformatics. 2009; 25: 1754-60

[49]

Price AL, Jones NC, Pevzner PA . De novo identification of repeat families in large genomes. Bioinformatics. 2005; 21 Suppl 1: i351-8

[50]

Xu Z, Wang H . LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 2007; 35: W265-8

[51]

Edgar RC, Myers EW . PILER: identification and classification of genomic repeats. Bioinformatics. 2005; 21 Suppl 1: i152-8

[52]

Bao W, Kojima KK, Kohany O . Repbase update, a database of repetitive elements in eukaryotic genomes. Mob DNA. 2015; 6: 11

[53]

Tarailo-Graovac M, Chen N . Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinformatics. 2009; Chapter 4: 4.10.1-14

[54]

Song X, Yang Q, Bai Y et al. Comprehensive analysis of SSRs and database construction using all complete gene-coding sequences in major horticultural and representative plants. Hortic Res. 2021; 8: 122

[55]

Song X, Li N, Guo Y et al. Comprehensive identification and characterization of simple sequence repeats based on the whole-genome sequences of 14 forest and fruit trees. Forestry Research. 2021; 1: 7

[56]

Nawrocki EP, Eddy SR . Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics. 2013; 29: 2933-5

[57]

Chan PP, Lowe TM . tRNAscan-SE: searching for tRNA genes in genomic sequences. Methods Mol Biol. 2019; 1962: 1-14

[58]

Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 1999; 27: 573-80

[59]

Korf I. Gene finding in novel genomes. BMC Bioinformatics. 2004; 5: 59

[60]

Stanke M, Morgenstern B . AUGUSTUS: a web server for gene prediction in eukaryotes that allows user-defined constraints. Nucleic Acids Res. 2005; 33: W465-7

[61]

Camacho C, Coulouris G, Avagyan V et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009; 10: 421

[62]

Birney E, Clamp M, Durbin R . GeneWise and Genomewise. Genome Res. 2004; 14: 988-95

[63]

Haas BJ, Salzberg SL, Zhu W et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 2008; 9: R7

[64]

Haas BJ, Delcher AL, Mount SM et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res. 2003; 31: 5654-66

[65]

Chen C, Chen H, Zhang Y et al. TBtools: an integrative toolkit developed for interactive analyses of big biological data. Mol Plant. 2020; 13: 1194-202

[66]

Emms DM, Kelly S . OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019; 20: 238

[67]

De Bie T, Cristianini N, Demuth JP et al. CAFE: a computational tool for the study of gene family evolution. Bioinformatics. 2006; 22: 1269-71

[68]

Edgar RC . MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004; 32: 1792-7

[69]

Stamatakis A . RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics. 2014; 30: 1312-3

[70]

Yang Z . PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol. 2007; 24: 1586-91

[71]

Kumar S, Stecher G, Suleski M et al. TimeTree: a resource for timelines, Timetrees, and divergence times. Mol Biol Evol. 2017; 34: 1812-9

[72]

Kim D, Langmead B, Salzberg SL . HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 2015; 12: 357-60

[73]

Trapnell C, Williams BA, Pertea G et al. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat Biotechnol. 2010; 28: 511-5

[74]

Anders S, Huber W . Differential expression analysis for sequence count data. Genome Biol. 2010; 11: R106

[75]

Wu T, Feng SY, Yang QH et al. Integration of the metabolome and transcriptome reveals the metabolites and genes related to nutritional and medicinal value in Coriandrum sativum. J Integr Agric. 2021; 20: 1807-18

[76]

Wang X, Shi X, Li Z et al. Statistical inference of chromosomal homology based on gene colinearity and applications to Arabidopsis and rice. BMC Bioinformatics. 2006; 7: 447

[77]

Sun P, Jiao B, Yang Y et al. WGDI: a user-friendly toolkit for evolutionary analyses of whole-genome duplications and ancestral karyotypes. Mol Plant. 2022; 15: 1841-51

[78]

Tang H, Bowers JE, Wang X et al. Synteny and collinearity in plant genomes. Science. 2008; 320: 486-8

[79]

Suyama M, Torrents D, Bork P . PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments. Nucleic Acids Res. 2006; 34: W609-12

[80]

Pei Q, Li N, Bai Y et al. Comparative analysis of the TCP gene family in celery, coriander and carrot (family Apiaceae). Vegetable Research. 2021; 1: 5

[81]

Pei Q, Yu T, Wu T et al. Comprehensive identification and analyses of the Hsf gene family in the whole-genome of three Apiaceae species. Hortic Plant J. 2021; 7: 457-68

[82]

Nakamura T, Yamada KD, Tomii K et al. Parallelization of MAFFT for large-scale multiple sequence alignments. Bioinformatics. 2018; 34: 2490-2

[83]

Price MN, Dehal PS, Arkin AP . FastTree: computing large minimum evolution trees with profiles instead of a distance matrix. Mol Biol Evol. 2009; 26: 1641-50

[84]

Yu T, Bai Y, Liu Z et al. Large-scale analyses of heat shock transcription factors and database construction based on whole-genome genes in horticultural and representative plants. Hortic Res. 2022; 9: uhac035

[85]

Wang Y, Song F, Zhu J et al. GSA: genome sequence archive. Genom Proteom Bioinform. 2017; 15: 14-8

[86]

BIG Data Center Members . Database resources of the BIG data center in 2019. Nucleic Acids Res. 2019; 47: D8-14

PDF (6667KB)

71

Accesses

0

Citation

Detail

Sections
Recommended

/