Genome assembly of Polygala tenuifolia provides insights into its karyotype evolution and triterpenoid saponin biosynthesis

Fanbo Meng , Tianzhe Chu , Pengmian Feng , Nan Li , Chi Song , Chunjin Li , Liang Leng , Xiaoming Song , Wei Chen

Horticulture Research ›› 2023, Vol. 10 ›› Issue (9) : 139

PDF (4647KB)
Horticulture Research ›› 2023, Vol. 10 ›› Issue (9) :139 DOI: 10.1093/hr/uhad139
Article
research-article
Genome assembly of Polygala tenuifolia provides insights into its karyotype evolution and triterpenoid saponin biosynthesis
Author information +
History +
PDF (4647KB)

Abstract

Polygala tenuifolia is a perennial medicinal plant that has been widely used in traditional Chinese medicine for treating mental diseases. However, the lack of genomic resources limits the insight into its evolutionary and biological characterization. In the present work, we reported the P. tenuifolia genome, the first genome assembly of the Polygalaceae family. We sequenced and assembled this genome by a combination of Illumnina, PacBio HiFi, and Hi-C mapping. The assembly includes 19 pseudochromosomes covering ∼92.68% of the assembled genome (∼769.62 Mb). There are 36 463 protein-coding genes annotated in this genome. Detailed comparative genome analysis revealed that P. tenuifolia experienced two rounds of whole genome duplication that occurred ∼39–44 and ∼18–20 million years ago, respectively. Accordingly, we systematically reconstructed ancestral chromosomes of P. tenuifolia and inferred its chromosome evolution trajectories from the common ancestor of core eudicots to the present species. Based on the transcriptomics data, enzyme genes and transcription factors involved in the synthesis of triterpenoid saponin in P. tenuifolia were identified. Further analysis demonstrated that whole-genome duplications and tandem duplications play critical roles in the expansion of P450 and UGT gene families, which contributed to the synthesis of triterpenoid saponins. The genome and transcriptome data will not only provide valuable resources for comparative and functional genomic researches on Polygalaceae, but also shed light on the synthesis of triterpenoid saponin.

Cite this article

Download citation ▾
Fanbo Meng, Tianzhe Chu, Pengmian Feng, Nan Li, Chi Song, Chunjin Li, Liang Leng, Xiaoming Song, Wei Chen. Genome assembly of Polygala tenuifolia provides insights into its karyotype evolution and triterpenoid saponin biosynthesis. Horticulture Research, 2023, 10 (9) : 139 DOI:10.1093/hr/uhad139

登录浏览全文

4963

注册一个新账户 忘记密码

Acknowledgements

This work was supported by the Natural Science Foundation of Sichuan (No. 2023NSFSC0683), Innovation Team and Talents Cultivation Program of the National Administration of Traditional Chinese Medicine (No: ZYYCXTD-D-202209), and the ‘Xinglin Scholar’ Discipline Talent Research Promotion Program of Chengdu University of TCM (No. MPRC2021036). The genome sequencing, Hi-C sequencing and primary assembly were performed with the help of Novogene.

Author contributions

W.C., L.L., and X.S. conceived the project and were responsible for the project initiation. W.C., L.L., and X.S. supervised and managed the project and research. Experiments and analyses were designed by W.C., L.L., and X.S. Data generation and bioinformatic analyses were performed by F.M., T.C., P. F., N.L., C.S., and C.L. The manuscript was organized, written and revised by F.M., X.S., L.L., and W.C. All authors read and revised the manuscript.

Data availability

The genome sequence and RNA-seq data of P. tenuifolia were deposited in the Genome Sequence Archive in BIG Data Center at Beijing Institute of Genomics (BIG, Chinese Academy of Sciences), under the accession numbers CRA009096 and CRA009246, which are publicly accessible at http://bigd.big.ac.cn/gsa. The annotations of the P. tenuifolia genome can be downloaded from the TCM Plant Genome Database (TCMPG: http://cbcb.cdutcm.edu.cn/TCMPG/resource/genomes/details/?id=TCMPG20196).

Conflict of interest statement

None declared.

Supplementary data

Supplementary data is available at Horticulture Research online.

References

[1]

Zhao X, Cui Y, Wu P et al. Polygalae radix: a review of its traditional uses, phytochemistry, pharmacology, toxicology, and pharmacokinetics. Fitoterapia. 2020; 147: 104759

[2]

Li H, Kim J, Tran HNK et al. Extract of Polygala tenuifolia, Angelica tenuissima, and Dimocarpus longan reduces behavioral defect and enhances autophagy in experimental models of Parkinson’s disease. NeuroMolecular Med. 2021; 23: 428-43

[3]

Vinh LB, Heo M, Phong NV et al. Bioactive compounds from Polygala tenuifolia and their inhibitory effects on lipopolysaccharide-stimulated pro-inflammatory cytokine production in bone marrow-derived dendritic cells. Plants (Basel). 2020; 9: 1240

[4]

Wang X, Li M, Cao Y et al. Tenuigenin inhibits LPS-induced inflammatory responses in microglia via activating the Nrf2-mediated HO-1 signaling pathway. Eur J Pharmacol. 2017; 809: 196-202

[5]

Li X, Zhao Y, Liu P et al. Senegenin inhibits hypoxia/Reoxygenation-induced neuronal apoptosis by upregulating RhoGDIα. Mol Neurobiol. 2015; 52: 1561-71

[6]

Zhu XQ, Li XM, Zhao YD et al. Effects of senegenin against hypoxia/reoxygenation-induced injury in PC12 cells. Chin J Integr Med. 2016; 22: 353-61

[7]

Jiang H, Liu T, Li L et al. Predicting the potential distribution of Polygala tenuifolia Willd. under climate change in China. PLoS One. 2016; 11: e0163718

[8]

Badouin H, Gouzy J, Grassa CJ et al. The sunflower genome provides insights into oil metabolism, flowering and Asterid evolution. Nature. 2017; 546: 148-52

[9]

Wang J, Sun P, Li Y et al. An overlooked Paleotetraploidization in Cucurbitaceae. Mol Biol Evol. 2018; 35: 16-26

[10]

Wang X, Jin D, Wang Z et al. Telomere-centric genome repatterning determines recurring chromosome number reductions during the evolution of eukaryotes. New Phytol. 2015; 205: 378-89

[11]

Wang Z, Li Y, Sun P et al. A high-quality Buxus austro-yunnanensis (Buxales) genome provides new insights into karyotype evolution in early eudicots. BMC Biol. 2022; 20: 216

[12]

Shen S, Li N, Wang Y et al. High-quality ice plant reference genome analysis provides insights into genome evolution and allows exploration of genes involved in the transition from C3 to CAM pathways. Plant Biotechnol J. 2022; 20: 2107-22

[13]

Song X, Wang J, Li N et al. Deciphering the high-quality genome sequence of coriander that causes controversial feelings. Plant Biotechnol J. 2020; 18: 1444-56

[14]

Jiang Z, Tu L, Yang W et al. The chromosome-level reference genome assembly for Panax notoginseng and insights into ginsenoside biosynthesis. Plant Commun. 2021; 2: 100113

[15]

Liao B, Shen X, Xiang L et al. Allele-aware chromosome-level genome assembly of Artemisia annua reveals the correlation between ADS expansion and artemisinin yield. Mol Plant. 2022; 15: 1310-28

[16]

Deng X, Zhao S, Liu X et al. Polygala tenuifolia: a source for anti-Alzheimer’s disease drugs. Pharm Biol. 2020; 58: 410-6

[17]

Wang Y, Zhang H, Ri HC et al. Deletion and tandem duplications of biosynthetic genes drive the diversity of triterpenoids in Aralia elata. Nat Commun. 2022; 13: 2224

[18]

Chung SY, Seki H, Fujisawa Y et al. A cellulose synthase-derived enzyme catalyses 3-O-glucuronosylation in saponin biosynthesis. Nat Commun. 2020; 11: 5664

[19]

Jin ML, Lee DY, Um Y et al. Isolation and characterization of an oxidosqualene cyclase gene encoding a β-amyrin synthase involved in Polygala tenuifolia Willd. saponin biosynthesis. Plant Cell Rep. 2014; 33: 511-9

[20]

Zhang FS, Zhang X, Wang QY et al. Cloning, Yeast Expression, and Characterization of a β-Amyrin C-28 Oxidase (CYP716A249) Involved in Triterpenoid Biosynthesis in Polygala tenuifolia. Biol Pharm Bull. 2020; 43: 1369-46

[21]

Young ND, Debellé F, Oldroyd GED et al. The Medicago genome provides insight into the evolution of rhizobial symbioses. Nature. 2011; 480: 520-4

[22]

Zhuang W, Chen H, Yang M et al. The genome of cultivated peanut provides insight into legume karyotypes, polyploid evolution and crop domestication. Nat Genet. 2019; 51: 865-76

[23]

Schmutz J, Cannon SB, Schlueter J et al. Genome sequence of the palaeopolyploid soybean. Nature. 2010; 463: 178-83

[24]

Lee DH, Cho WB, Park B et al. The complete chloroplast genome of Polygala tenuifolia, a critically endangered species in Korea. Mitochondrial DNA Part B. 2020; 5: 1919-20

[25]

Zuo Y, Mao Y, Shang S et al. The complete chloroplast genome of Polygala japonica Houtt. (Polygalaceae), a medicinal plant in China. Mitochondrial DNA Part B. 2021; 6: 239-40

[26]

Ma J, Wang J, Li C et al. The complete chloroplast genome characteristics of Polygala crotalarioides Buch.-ham. ex DC. (Polygalaceae) from Yunnan, China. Mitochondrial DNA Part B. 2021; 6: 2838-40

[27]

Wu S, Han B, Jiao Y . Genetic contribution of Paleopolyploidy to adaptive evolution in angiosperms. Mol Plant. 2020; 13: 59-71

[28]

Zhang L, Wu S, Chang X et al. The ancient wave of polyploidization events in flowering plants and their facilitated adaptation to environmental stress. Plant Cell Environ. 2020; 43: 2847-56

[29]

Soltis PS, Marchant DB, Van de Peer Y et al. Polyploidy and genome evolution in plants. Curr Opin Genet Dev. 2015; 35: 119-25

[30]

Freeling M, Scanlon MJ, Fowler JE . Fractionation and subfunctionalization following genome duplications: mechanisms that drive gene content and their consequences. Curr Opin Genet Dev. 2015; 35: 110-8

[31]

Cheng F, Wu J, Cai X et al. Gene retention, fractionation and subgenome differences in polyploid plants. Nat Plants. 2018; 4: 258-68

[32]

Qiao X, Li Q, Yin H et al. Gene duplication and evolution in recurring polyploidization-diploidization cycles in plants. Genome Biol. 2019; 20: 38

[33]

Thimmappa R, Geisler K, Louveau T et al. Triterpene biosynthesis in plants. Annu Rev Plant Biol. 2014; 65: 225-57

[34]

Nützmann HW, Osbourn A . Gene clustering in plant specialized metabolism. Curr Opin Biotechnol. 2014; 26: 91-9

[35]

Nützmann HW, Huang A, Osbourn A . Plant metabolic clusters - from genetics to genomics. New Phytol. 2016; 211: 771-89

[36]

Kautsar SA, Suarez Duran HG, Blin K et al. plantiSMASH: automated identification, annotation and expression analysis of plant biosynthetic gene clusters. Nucleic Acids Res. 2017; 45: w55-63

[37]

Belton JM, McCord RP, Gibcus JH et al. Hi-C: a comprehensive technique to capture the conformation of genomes. Methods. 2012; 58: 268-76

[38]

Cheng H, Concepcion GT, Feng X et al. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021; 18: 170-5

[39]

Zhang X, Zhang S, Zhao Q et al. Assembly of allele-aware, chromosomal-scale autopolyploid genomes based on hi-C data. Nat Plants. 2019; 5: 833-45

[40]

Wingett S, Ewels P, Furlan-Magaril M et al. HiCUP: pipeline for mapping and processing hi-C data. F1000Res. 2015; 4: 1310

[41]

Durand NC, Shamim MS, Machol I et al. Juicer provides a one-click system for analyzing loop-resolution hi-C experiments. Cell Syst. 2016; 3: 95-8

[42]

Manni M, Berkeley MR, Seppey M et al. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol. 2021; 38: 4647-54

[43]

Parra G, Bradnam K, Korf I . CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics. 2007; 23: 1061-7

[44]

Li H, Durbin R . Fast and accurate short read alignment with burrows-wheeler transform. Bioinformatics. 2009; 25: 1754-60

[45]

Jurka J, Kapitonov VV, Pavlicek A et al. Repbase update, a database of eukaryotic repetitive elements. Cytogenet Genome Res. 2005; 110: 462-7

[46]

Xu Z, Wang H . LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 2007; 35: w265-8

[47]

Price AL, Jones NC, Pevzner PA . De novo identification of repeat families in large genomes. Bioinformatics. 2005; 21: i351-8

[48]

Stanke M, Waack S . Gene prediction with a hidden Markov model and a new intron submodel. Bioinformatics. 2003; 19: ii215-25

[49]

Burge C, Karlin S . Prediction of complete gene structures in human genomic DNA. J Mol Biol. 1997; 268: 78-94

[50]

Majoros WH, Pertea M, Salzberg SL . TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics. 2004; 20: 2878-9

[51]

Korf I . Gene finding in novel genomes. BMC Bioinformatics. 2004; 5: 59

[52]

Birney E, Clamp M, Durbin R . GeneWise and Genomewise. Genome Res. 2004; 14: 988-95

[53]

Grabherr MG, Haas BJ, Yassour M et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol. 2011; 29: 644-52

[54]

Trapnell C, Pachter L, Salzberg SL . TopHat: discovering splice junctions with RNA-Seq. Bioinformatics. 2009; 25: 1105-11

[55]

Trapnell C, Williams BA, Pertea G et al. Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation. Nat Biotechnol. 2010; 28: 511-5

[56]

Haas BJ, Delcher AL, Mount SM et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res. 2003; 31: 5654-66

[57]

Bairoch A, Apweiler R . The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000. Nucleic Acids Res. 2000; 28: 45-8

[58]

Mulder N, Apweiler R . InterPro and InterProScan: tools for protein sequence classification and comparison. Methods Mol Biol. 2007; 396: 59-70

[59]

Kanehisa M, Goto S . KEGG: Kyoto encyclopedia of genes and genomes. Nucleic Acids Res. 2000; 28: 27-30

[60]

Lowe TM, Eddy SR . tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res. 1997; 25: 955-64

[61]

Griffiths-Jones S, Moxon S, Marshall M et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 2004; 33: D121-4

[62]

Chen C, Chen H, Zhang Y et al. TBtools: an integrative toolkit developed for interactive analyses of big biological data. Mol Plant. 2020; 13: 1194-202

[63]

Chen C, Wu Y, Xia R . A painless way to customize Circos plot: from data preparation to visualization using TBtools. iMeta. 2022; 1: e35

[64]

Li L, Stoeckert CJ Jr, Roos DS . OrthoMCL: identification of ortholog groups for eukaryotic genomes. Genome Res. 2003; 13: 2178-89

[65]

Edgar RC . MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004; 32: 1792-7

[66]

Stamatakis A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics. 2014; 30: 1312-3

[67]

Yang Z . PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol. 2007; 24: 1586-91

[68]

De Bie T, Cristianini N, Demuth JP et al. CAFE: a computational tool for the study of gene family evolution. Bioinformatics. 2006; 22: 1269-71

[69]

Camacho C, Coulouris G, Avagyan V et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009; 10: 421

[70]

Sun P, Jiao B, Yang Y et al. WGDI: a user-friendly toolkit for evolutionary analyses of whole-genome duplications and ancestral karyotypes. Mol Plant. 2022; 15: 1841-51

[71]

Wang X, Shi X, Li Z et al. Statistical inference of chromosomal homology based on gene colinearity and applications to Arabidopsis and rice. BMC Bioinformatics. 2006; 7: 447

[72]

Wang J, Yuan J, Yu J et al. Recursive paleohexaploidization shaped the durian genome. Plant Physiol. 2019; 179: 209-19

[73]

Yang Y, Sun P, Lv L et al. Prickly waterlily and rigid hornwort genomes shed light on early angiosperm evolution. Nat Plants. 2020; 6: 215-22

[74]

Chanderbali AS, Jin L, Xu Q et al. Buxus and Tetracentron genomes help resolve eudicot genome history. Nat Commun. 2022; 13: 643

[75]

Kim D, Paggi JM, Park C et al. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol. 2019; 37: 907-15

[76]

Pertea M, Pertea GM, Antonescu CM et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol. 2015; 33: 290-5

[77]

Liao Y, Smyth GK, Shi W . featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014; 30: 923- 30

[78]

Love MI, Huber W, Anders S . Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014; 15: 550

[79]

Wu T, Hu E, Xu S et al. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data. Innovation (Camb). 2021; 2: 100141

[80]

Potter SC, Luciani A, Eddy SR et al. HMMER web server: 2018 update. Nucleic Acids Res. 2018; 46: w200-4

[81]

Meng F, Tang Q, Chu T et al. TCMPG: an integrative database for traditional Chinese medicine plant genomes. Hortic Res. 2022; 9: uhac060

[82]

Kumar S, Stecher G, Li M et al. MEGA X: molecular evolutionary genetics analysis across computing platforms. Mol Biol Evol. 2018; 35: 1547-9

[83]

Letunic I, Bork P . Interactive tree of life (iTOL) v5: an online tool for phylogenetic tree display and annotation. Nucleic Acids Res. 2021; 49: w293-6

[84]

Jin J, Tian F, Yang DC et al. PlantTFDB 4.0: toward a central hub for transcription factors and regulatory interactions in plants. Nucleic Acids Res. 2017; 45: d1040-5

[85]

Shannon P, Markiel A, Ozier O et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003; 13: 2498-504

[86]

Wang Y, Tang H, DeBarry JD et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res. 2012; 40: e49

PDF (4647KB)

93

Accesses

0

Citation

Detail

Sections
Recommended

/