Chromosomal-level genome and multi-omics dataset of Pueraria lobata var. thomsonii provide new insights into legume family and the isoflavone and puerarin biosynthesis pathways

Xiaohong Shang , Xinxin Yi , Liang Xiao , Yansheng Zhang , Ding Huang , Zhengbao Xia , Kunpeng Ou , Ruhong Ming , Wendan Zeng , Dongqing Wu , Sheng Cao , Liuyin Lu , Huabing Yan

Horticulture Research ›› 2022, Vol. 9 ›› Issue (1) : uhab035

PDF (3374KB)
Horticulture Research ›› 2022, Vol. 9 ›› Issue (1) :uhab035 DOI: 10.1093/hr/uhab035
Article
research-article
Chromosomal-level genome and multi-omics dataset of Pueraria lobata var. thomsonii provide new insights into legume family and the isoflavone and puerarin biosynthesis pathways
Author information +
History +
PDF (3374KB)

Abstract

Pueraria lobata var. thomsonii (hereinafter abbreviated as Podalirius thomsonii), a member of the legume family, is one of the important traditional Chinese herbal medicines, and its puerarin extract is widely used in the health and pharmaceutical industry. Here, we assembled a high-quality genome of P. thomsonii using long-read single-molecule sequencing and Hi-C technologies. The genome assembly is ∼1.37 Gb in size and consists of 5145 contigs with a contig N50 of 593.70 kb, further clustered into 11 pseudochromosomes. Genome structural annotation resulted in ∼869.33 Mb (∼62.70% of the genome) repeat regions and 45 270 protein-coding genes. Genome evolution analysis revealed that P. thomsonii is most closely related to soybean and underwent two ancient whole-genome duplication events; one was in the common ancestor shared by legume species and the other occurred independently at around 7.2 million years ago, after its speciation. A total of 2373 gene families were found to be unique in P. thomsonii compared with five other legume species. Genes and metabolites related to puerarin content in tuberous tissues were characterized. A total of 572 genes that were upregulated in the puerarin biosynthesis pathway were identified, and 235 candidate genes were further enriched by omics data. Furthermore, we identified six 8-C-glucosyltransferase (8-C-GT) candidate genes significantly involved in puerarin metabolism. Our study filled a key genomic gap in the legume family, and provided valuable multi-omic resources for the genetic improvement of P. thomsonii .

Cite this article

Download citation ▾
Xiaohong Shang, Xinxin Yi, Liang Xiao, Yansheng Zhang, Ding Huang, Zhengbao Xia, Kunpeng Ou, Ruhong Ming, Wendan Zeng, Dongqing Wu, Sheng Cao, Liuyin Lu, Huabing Yan. Chromosomal-level genome and multi-omics dataset of Pueraria lobata var. thomsonii provide new insights into legume family and the isoflavone and puerarin biosynthesis pathways. Horticulture Research, 2022, 9 (1) : uhab035 DOI:10.1093/hr/uhab035

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Wang S, Zhang S, Wang S et al. A comprehensive review on Pueraria: insights on its chemistry and medicinal value. Biomed Pharmacother 2020; 131: 110734.

[2]

Zhao Z, Guo P, Brand E . A concise classification of bencao (materia medica). Chin Med. 2018; 13: 18.

[3]

Egan AN . Economic and ethnobotanical uses of tubers in the genus Pueraria DC. Legume 2020; 19: 24.

[4]

WFO. Pueraria montana var. chinensis (Ohwi) Sanjappa & Pradeep . Date accessed 21 May 2021. http://www.worldfloraonline.org/taxon/wfo-0000193855. 2021.

[5]

Li S, Xu L, Chen D et al. Fabaceae (Leguminosae). Flora of China. 2010; 41: 226.

[6]

Heider B, Fischer E, Berndl T et al. Analysis of genetic variation among accessions of Pueraria montana (Lour.) Merr. var. lobata and Pueraria phaseoloides (Roxb.) Benth. based on RAPD markers . Genet Resour Crop Evol. 2007; 54: 529-42.

[7]

Jewett D, Jiang C, Britton K et al. Characterizing specimens of kudzu and related taxa with RAPD’s. Castanea. 2003; 68: 254-60.

[8]

Haynsen MS, Vatanparast M, Mahadwar G et al. De novo transcriptome assembly of Pueraria montana var. lobata and Neustanthus phaseoloides for the development of eSSR and SNP markers: narrowing the US origin(s) of the invasive kudzu . BMC Genomics. 2018; 19: 439.

[9]

Zhang G, Liu J, Gao M et al. Tracing the edible and medicinal plant Pueraria montana and its products in the marketplace yields subspecies level distinction using DNA barcoding and DNA metabarcoding. Front Pharmacol 2020; 11: 336.

[10]

Zhou YX, Zhang H, Peng C . Puerarin: a review of pharmacological effects. Phytother Res 2014; 28: 961-75.

[11]

Chen YG, Song YL, Wang Y et al. Metabolic differentiations of Pueraria lobata and Pueraria thomsonii using (1)H NMR spectroscopy and multivariate statistical analysis . J Pharm Biomed Anal. 2014; 93: 51-8.

[12]

Chen SB, Liu HP, Tian RT et al. High-performance thin-layer chromatographic fingerprints of isoflavonoids for distinguishing between Radix Puerariae lobate and Radix Puerariae Thomsonii . J Chromatogr A. 2006; 1121: 114-9.

[13]

Wang X, Li S, Li J et al. De novo transcriptome sequencing in Pueraria lobata to identify putative genes involved in isoflavones biosynthesis. Plant Cell Rep. 2015; 34: 733-43.

[14]

He X, Blount JW, Ge S et al. A genomic approach to isoflavone biosynthesis in kudzu (Pueraria lobata) . Planta. 2011; 233: 843-55.

[15]

Wang X, Li C, Zhou Z et al. Identification of three (iso)flavonoid glucosyltransferases from Pueraria lobata . Front Plant Sci 2019; 10: 28.

[16]

Wang X, Li C, Zhou C et al. Molecular characterization of the C-glucosylation for puerarin biosynthesis in Pueraria lobata . Plant J 2017; 90: 535-46.

[17]

Kreplav J, Madoui MA, Capal P et al. A reference genome for pea provides insight into legume genome evolution. Nat Genet. 2019; 51: 1411-22.

[18]

Vanneste K, Baele G, Maere S et al. Analysis of 41 plant genomes supports a wave of successful genome duplications in association with the cretaceous-Paleogene boundary. Genome Res. 2014; 24: 1334-47.

[19]

Schmutz J, Cannon SB, Schlueter J et al. Genome sequence of the palaeopolyploid soybean. Nature 2010; 463: 178-83.

[20]

Han R, Takahashi H, Nakamura M et al. Transcriptomic landscape of Pueraria lobata demonstrates potential for phytochemical study. Front Plant Sci 2015; 6: 426.

[21]

Doyle JJ, Doyle JL . A rapid DNA isolation procedure for small quantities of fresh leaf tissue. Phytochem Bull. 1987; 19: 11-5.

[22]

Belton JM, McCord RP, Gibcus J et al. Hi-C: a comprehensive technique to capture the conformation of genomes. Methods. 2012; 58: 268-76.

[23]

Yang X, Liu D, Wu J et al. HTQC: a fast quality control toolkit for Illumina sequencing data. BMC Bioinformatics. 2013; 14: 33.

[24]

Marcais G, Kingsford C . A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 2011; 27: 764-70.

[25]

Chin CS, Peluso P, Sedlazeck FJ et al. Phased diploid genome assembly with single-molecule real-time sequencing. Nat Methods. 2016; 13: 1050-4.

[26]

Vaser R, Sovic I, Nagarajan N et al. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 2017; 27: 737-46.

[27]

Hu J, Fan J, Sun Z et al. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics. 2020; 36: 2253-5.

[28]

Chen S, Zhou Y, Chen Y et al. Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics. 2018; 34: i884-90.

[29]

Servant N, Varoquaux N, Lajoie BR et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 2015; 16: 259.

[30]

Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM[J]. 2013; 1-3. arXiv preprint arXiv:1303.3997.

[31]

Xu Z, Wang H . LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 2007; 35: W265-8.

[32]

Price AL, Jones NC, De Pevzner PA . De novo identification of repeat families in large genomes. Bioinformatics. 2005; 21: i351-8.

[33]

Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 1999; 27: 573-80.

[34]

Tarailo-Graovac M, Chen N . Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinformatics 2009; 25: 4.10.1-4.10.14.

[35]

Cantarel BL, Korf I, Robb SMC et al. MAKER: an easy-to-use annotation pipeline designed for emerging model organism genomes. Genome Res. 2007; 18: 188-96.

[36]

Boeckmann B, Bairoch A, Apweiler R et al. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Res. 2003; 31: 365-70.

[37]

Altschul SF, Madden TL, Schäffer AA et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997; 25: 3389-402.

[38]

Camacho C, Coulouris G, Avagyan V et al. BLAST+: architecture and applications. BMC Bioinformatics. 2009; 10: 421.

[39]

Kanehisa M, Goto S, Sato Y et al. KEGG for integration and interpretation of large-scale molecular data sets. Nucleic Acids Res. 2012; 40: D109-14.

[40]

Ashburner M, Ball CA, Blake JA et al. Gene ontology: tool for the unification of biology. Nature Genet. 2000; 25: 25-9.

[41]

Conesa A, Gotz S . Blast2GO: a comprehensive suite for functional analysis in plant genomics. Int J Plant Genomics. 2008; 2008: 619832.

[42]

Li L, Stoeckert CJ Jr, Roos DS . OrthoMCL: identification of ortholog groups for eukaryotic genomes. Genome Res. 2003; 13: 2178-89.

[43]

Edgar RC . MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004; 32: 1792-7.

[44]

Stamatakis A . RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics. 2014; 30: 1312-3.

[45]

Han MV, Thomas GW, Lugo-Martinez J et al. Estimating gene gain and loss rates in the presence of error in genome assembly and annotation using CAFE 3. Mol Biol Evol. 2013; 30: 1987-97.

[46]

Yang Z . PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol. 2007; 24: 1586-91.

[47]

Tang H, Bowers JE, Wang X et al. Synteny and collinearity in plant genomes. Science. 2008; 320: 486-8.

[48]

Qing Z, Liu J, Yi X et al. The chromosome-level Hemerocallis citrina Borani genome provides new insights into the rutin biosynthesis and the lack of colchicine. Hortic Res. 2021; 8: 89.

[49]

Wang D, Zhang Y, Zhang Z et al. KaKs_Calculator 2.0: a toolkit incorporating gamma-series methods and sliding window strategies. Genomics Proteomics Bioinformatics. 2010; 8: 77-80.

[50]

Kim D, Langmead B, Salzberg SL . HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 2015; 12: 357-60.

[51]

Pertea M, Pertea GM, Antonescu CM et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol. 2015; 33: 290-5.

[52]

Langmead B, Salzberg SL . Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012; 9: 357-9.

[53]

Li B, Dewey CN . RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics. 2011; 12: 323.

[54]

Love MI, Huber W, Anders S . Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014; 15: 550.

[55]

Chen W, Gong L, Guo Z et al. A novel integrated method for large-scale detection, identification, and quantification of widely targeted metabolites: application in the study of rice metabolomics. Mol Plant. 2013; 6: 1769-80.

[56]

Chen Y, Zhang R, Song Y et al. RRLC-MS/MS-based metabolomics combined with in-depth analysis of metabolic correlation network: finding potential biomarkers for breast cancer. Analyst. 2009; 134: 2003-11.

[57]

Kolde R, Kolde MR . Package ‘pheatmap’. R package. 2015; 1: 790.

[58]

Chong J, Xia J . MetaboAnalystR: an R package for flexible and reproducible analysis of metabolomics data. Bioinformatics. 2018; 34: 4313-4.

[59]

Finn RD, Bateman A, Clements J et al. Pfam: the protein families database. Nucleic Acids Res. 2014; 42: D222-30.

[60]

Hunter S, Apweiler R, Attwood TK et al. InterPro: the integrative protein signature database. Nucleic Acids Res. 2009; 37: D211-5.

[61]

Finn RD, Clements J, Eddy SR . HMMER web server: interactive sequence similarity searching. Nucleic Acids Res. 2011; 39: W29-37.

PDF (3374KB)

45

Accesses

0

Citation

Detail

Sections
Recommended

/