Pangenome of water caltrop reveals structural variations and asymmetric subgenome divergence after allopolyploidization

Xinyi Zhang , Yang Chen , Lingyun Wang , Ye Yuan , Mingya Fang , Lin Shi , Ruisen Lu , Hans Peter Comes , Yazhen Ma , Yuanyuan Chen , Guizhou Huang , Yongfeng Zhou , Zhaisheng Zheng , Yingxiong Qiu

Horticulture Research ›› 2023, Vol. 10 ›› Issue (11) : 203

PDF (4269KB)
Horticulture Research ›› 2023, Vol. 10 ›› Issue (11) :203 DOI: 10.1093/hr/uhad203
Article
research-article
Pangenome of water caltrop reveals structural variations and asymmetric subgenome divergence after allopolyploidization
Author information +
History +
PDF (4269KB)

Abstract

Water caltrop (Trapa spp., Lythraceae) is a traditional but currently underutilized non-cereal crop. Here, we generated chromosome-level genome assemblies for the two diploid progenitors of allotetraploid Trapa natans (4x, AABB), i.e., diploid T. natans (2x, AA) and Trapa incisa (2x, BB). In conjunction with four published (sub)genomes of Trapa, we used gene-based and graph-based pangenomic approaches and a pangenomic transposable element (TE) library to develop Trapa genomic resources. The pangenome displayed substantial gene-content variation with dispensable and private gene clusters occupying a large proportion (51.95%) of the total cluster sets in the six (sub)genomes. Genotyping of presence-absence variation (PAVs) identified 40 453 PAVs associated with 2570 genes specific to A- or B-lineages, of which 1428 were differentially expressed, and were enriched in organ development process, organic substance metabolic process and response to stimulus. Comparative genome analyses showed that the allotetraploid T. natans underwent asymmetric subgenome divergence, with the B-subgenome being more dominant than the A-subgenome. Multiple factors, including PAVs, asymmetrical amplification of TEs, homeologous exchanges (HEs), and homeolog expression divergence, together affected genome evolution after polyploidization. Overall, this study sheds lights on the genome architecture and evolution of Trapa, and facilitates its functional genomic studies and breeding program.

Cite this article

Download citation ▾
Xinyi Zhang, Yang Chen, Lingyun Wang, Ye Yuan, Mingya Fang, Lin Shi, Ruisen Lu, Hans Peter Comes, Yazhen Ma, Yuanyuan Chen, Guizhou Huang, Yongfeng Zhou, Zhaisheng Zheng, Yingxiong Qiu. Pangenome of water caltrop reveals structural variations and asymmetric subgenome divergence after allopolyploidization. Horticulture Research, 2023, 10 (11) : 203 DOI:10.1093/hr/uhad203

登录浏览全文

4963

注册一个新账户 忘记密码

Acknowledgements

We thank Lin Cheng (Agricultural Genomics Institute at Shenzhen) and Shuo Cao (Agricultural Genomics Institute at Shenzhen) for their assistance with data analyses; Emmanuel Nyongesa Waswa (Wuhan Botanical Garden) for his assistance with polishing the manuscript; and two anonymous referees for valuable advice and comments that have substantially improved the manuscript. This work was supported by the collaborative program of Chinese Academy of Agricultural Sciences (CAAS)- Jinhua Academy of Agricultural Sciences, funded by Jinhua City of Zhejiang Province, and the Research Grant from Wuhan Botanic Garden (E1559901)

Author contributions

Y.Q. and Z.Z. conceived the project. X.Z., Y.C., and Y.Q. performed experiments and coordinated research activities. X.Z., Y.C., L.W., Y.Y., L.S., and M.F. collected the samples. X.Z. analysed data. X.Z. and Y.Q. drafted the manuscript. Y.Q., H.P.C., Y.Z., and Y.M. revised and finalized the manuscript. All authors read and approved the manuscript. X.Z. and Y.C. contributed equally to this work.

Data availability

The whole genome sequencing data for diploid T. natans have been deposited under NCBI BioProject PRJNA932942 and GSA BioProject PRJCA016421. The whole genome sequencing data for diploid T. incisa have been deposited under NCBI BioProject PRJNA933001 and GSA BioProject PRJCA016421. The transcriptome sequencing data of diploid T. incisa and allotetraploid T. natans have been deposited under NCBI BioProject PRJNA941110 and PRJNA731291, respectively. The re-sequencing (57 accessions) and transcriptome sequencing data of diploid Trapa natans for genome annotation were obtained from the NCBI BioProject PRJNA725399 of Lu et al. [10].

Conflict of interests

All authors confirm that they have no conflict of interest.

Supplementary data

Supplementary data is available at Horticulture Research online.

References

[1]

Jain S, Dutta GS, eds. Neglected and Underutilized Crops - Towards Nutritional Security and Sustainability, Biotechnology of Neglected and Underutilized Crops. Dordrecht: Springer; 2012: p. v.

[2]

Chang Y, Liu H, Liu M, et al. The draft genomes of five agriculturally important African orphan crops. GigaScience. 2019; 8: giy152

[3]

Li X, Yadav R, Siddique KHM . Neglected and underutilized crop species: the key to improving dietary diversity and fighting hunger and malnutrition in Asia and the Pacific. Front Nutr. 2020; 7: 593711

[4]

Ye C, Fan L . Orphan crops and their wild relatives in the genomic era. Mol Plant. 2021; 14: 27-39

[5]

Dawson IK, Powell W, Hendre P, et al. The role of genetics in mainstreaming the production of new and orphan crops to diversify food systems and support human nutrition. New Phytol. 2019; 224: 37-54

[6]

Takano A, Kadono Y . Allozyme variations and classification of Trapa (Trapaceae) in Japan . Aquat Bot. 2005; 83: 108-18

[7]

Ding B, Jin X . Taxonomic notes on genus Trapa L. (Trapaceae) in China . Guihaia. 2015; 40: 1-15

[8]

Hoque A, Davey MR, Arima S . Water chestnut: potential of biotechnology for crop improvement. J New Seeds. 2009; 10: 180-95

[9]

Guo Y, Wu R, Sun G, et al. Neolithic cultivation of water chestnuts (Trapa L.) at Tianluoshan (7000-6300 cal BP), Zhejiang Province, China. Sci Rep. 2017; 7: 16206

[10]

Lu R, Chen Y, Zhang X, et al. Genome sequencing and transcriptome analyses provide insights into the origin and domestication of water caltrop (Trapa spp., Lythraceae). Plant Biotechnol J. 2022; 20: 761-76

[11]

Qu M, Fan X, Hao C, et al. Chromosome-level assemblies of cultivated water chestnut Trapa bicornis and its wild relative Trapa incisa. Scientific Data. 2023; 10: 407

[12]

Gaut BS, Seymour DK, Liu Q, et al. Demography and its effects on genomic variation in crop domestication. Nature Plants. 2018; 4: 512-20

[13]

Hämälä T, Wafula EK, Guiltinan MJ, et al. Genomic structural variants constrain and facilitate adaptation in natural populations of Theobroma cacao, the chocolate tree . Proc Natl Acad Sci. 2021; 118: e2102914118

[14]

Kou Y, Liao Y, Toivainen T, et al. Evolutionary genomics of structural variation in asian rice (Oryza sativa) domestication . Mol Biol Evol. 2020; 37: 3507-24

[15]

Della Coletta R, Qiu Y, Ou S, et al. How the pan-genome is changing crop genomics and improvement. Genome Biol. 2021; 22: 3

[16]

Danilevicz MF, Tay Fernandez CG, Marsh JI, et al. Plant pangenomics: approaches, applications and advancements. Curr Opin Plant Biol. 2020; 54: 18-25

[17]

Torkamaneh D, Lemay M-A, Belzile F . The pan-genome of the cultivated soybean (PanSoy) reveals an extraordinarily conserved gene content. Plant Biotechnol J. 2021; 19: 1852-62

[18]

Liu Y, Du H, Li P, et al. Pan-genome of wild and cultivated soybeans. Cell. 2020; 182: 162-176.e13

[19]

Qin P, Lu H, Du H, et al. Pan-genome analysis of 33 genetically diverse rice accessions reveals hidden genomic variations. Cell. 2021; 184: 3542-3558.e16

[20]

Zhao Q, Feng Q, Lu H, et al. Pan-genome analysis highlights the extent of genomic variation in cultivated and wild rice. Nat Genet. 2018; 50: 278-84

[21]

Gui S, Wei W, Jiang C, et al. A pan-zea genome map for enhancing maize improvement. Genome Biol. 2022; 23: 178

[22]

Li L, Zhang Z, Wang Z, et al. Genome sequences of five Sitopsis species of Aegilops and the origin of polyploid wheat B subgenome. Mol Plant. 2022; 15: 488-503

[23]

Tao Y, Luo H, Xu J, et al. Extensive variation within the pan-genome of cultivated and wild sorghum. Nature Plants. 2021; 7: 766-73

[24]

Yu J, Golicz AA, Lu K, et al. Insight into the evolution and functional characteristics of the pan-genome assembly from sesame landraces and modern cultivars. Plant Biotechnol J. 2019; 17: 881-92

[25]

Zhao J, Bayer PE, Ruperao P, et al. Trait associations in the pangenome of pigeon pea (Cajanus cajan). Plant Biotechnol J. 2020; 18: 1946-54

[26]

Catlin NS, Josephs EB . The important contribution of transposable elements to phenotypic variation and evolution. Curr Opin Plant Biol. 2022; 65: 102140

[27]

Ou S, Collins T, Qiu Y, et al. Differences in activity and stability drive transposable element variation in tropical and temperate maize. bioRxiv. 2022; 2022.10.09.511471, preprint: not peer reviewed.

[28]

Mérot C, Oomen RA, Tigano A, et al. A roadmap for understanding the evolutionary significance of structural genomic variation. Trends Ecol Evol. 2020; 35: 561-72

[29]

Graham SA, Hall J, Sytsma K, et al. Phylogenetic analysis of the Lythraceae based on four gene regions and morphology. Int J Plant Sci. 166: 995-1017

[30]

Berger BA, Kriebel R, Spalink D, et al. Divergence times, historical biogeography, and shifts in speciation rates of Myrtales. Mol Phylogenet Evol. 2016; 95: 116-36

[31]

Chalhoub B, Denoeud F, Liu S, et al. Early allopolyploid evolution in the post-Neolithic Brassica napus oilseed genome . Science. 2014; 345: 950-3

[32]

Alfasane MA, Khondker M, Rahman MM . Biochemical composition of the fruits of water chestnut (Trapa bispinosa Roxb.). Dhaka Univ J Biol Sci. 2011; 20: 95-8

[33]

Subbahmanyan V, Rama RG, Kuppuswamy S, et al. Nutritive value of water chestnut (Singhara). Bull Cent Food Tech Res Inst. 1954; 3: 134-5

[34]

Hurgobin B, Golicz AA, Bayer PE, et al. Homoeologous exchange is a major cause of gene presence/absence variation in the amphidiploid Brassica napus. Plant Biotechnol J. 2018; 16: 1265-1274.

[35]

Bayer PE, Scheben A, Golicz AA, et al. Modelling of gene loss propensity in the pangenomes of three Brassica species suggests different mechanisms between polyploids and diploids. Plant Biotechnol J. 2021; 19: 2488-2500.

[36]

Li A, Wang J, Sun K, et al. Two reference-quality sea snake genomes reveal their divergent evolution of adaptive traits and venom systems. Mol Biol Evol. 2021; 38: 4867-83

[37]

Zhou Y, Minio A, Massonnet M, et al. The population genetics of structural variants in grapevine domestication. Nature Plants. 2019; 5: 965-79

[38]

Ding B, Jin X . Taxonomic notes on genus Trapa L. (Trapaceae) in China . Guihaia 2019; 40: 1-15.

[39]

Sirén J, Monlong J, Chang X, et al. Pangenomics enables genotyping of known structural variants in 5202 diverse genomes. Science. 2021; 374: abg8871

[40]

Crysnanto D, Pausch H . Bovine breed-specific augmented reference graphs facilitate accurate sequence read mapping and unbiased variant discovery. Genome Biol. 2020; 21: 184

[41]

Zhou Y, Zhang Z, Bao Z, et al. Graph pangenome captures missing heritability and empowers tomato breeding. Nature. 2022; 606: 527-34

[42]

Zumajo-Cardona C, Aguirre M, Castillo-Bravo R, et al. Maternal control of triploid seed development by the TRANSPARENT TESTA 8 (TT8) transcription factor in Arabidopsis thaliana. Sci Rep. 2023; 13: 1316

[43]

Bird KA, VanBuren R, Puzey JR, et al. The causes and consequences of subgenome dominance in hybrids and recent polyploids. New Phytol. 2018; 220: 87-93

[44]

Sun Y, Liu Y, Shi J, et al. Biased mutations and gene losses underlying diploidization of the tetraploid broomcorn millet genome. Plant J. 2023; 113: 787-801

[45]

Marçais G, Kingsford C . A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 2011; 27: 764-70

[46]

Ranallo-Benavidez TR, Jaron KS, Schatz MC . GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat Commun. 2020; 11: 1432

[47]

Cheng H, Concepcion GT, Feng X, et al. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 2021; 18: 170-5

[48]

Vasimuddin M, Misra S, Li H, et al. Efficient architecture-aware acceleration of BWA-MEM for multicore systems. IEEE International Parallel and Distributed Processing Symposium. 2019; 314-24

[49]

Zhang X, Zhang S, Zhao Q, et al. Assembly of allele-aware, chromosomal-scale autopolyploid genomes based on Hi-C data. Nat Plants. 2019; 5: 833-845

[50]

Durand NC, Robinson JT, Shamim MS, et al. Juicebox provides a visualization system for hi-C contact maps with unlimited zoom. Cell Systems. 2016; 3: 99-101

[51]

Simão FA, Waterhouse RM, Ioannidis P, et al. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 2015; 31: 3210-2

[52]

Parra G, Bradnam K, Korf I . CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics. 2007; 23: 1061-7

[53]

Koren S, Walenz BP, Berlin K, et al. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 2017; 27: 722-36

[54]

Ou S, Su W, Liao Y, et al. Benchmarking transposable element annotation methods for creation of a streamlined, comprehensive pipeline. Genome Biol. 2019; 20: 275

[55]

Yan H, Bombarely A, Li S . DeepTE: a computational method for de novo classification of transposons with convolutional neural network. Bioinformatics. 2020; 36: 4269-75

[56]

Stanke M, Keller O, Gunduz I, et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic Acids Res. 2006; 34: W435-9

[57]

Majoros WH, Pertea M, Salzberg SL . TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics. 2004; 20: 2878-9

[58]

Burge C, Karlin S . Prediction of complete gene structures in human genomic. J Mol Biol. 1997; 268: 78-94

[59]

Korf I. Gene finding in novel genomes. BMC Bioinformatics. 2004; 5: 59

[60]

Blanco E, Parra G, Guigó R . Using geneid to identify genes. Curr Protoc Bioinformatics. 2007; 18: 4.3.1-4.3.28

[61]

Healey AL, Shepherd M, King GJ, et al. Pests, diseases, and aridity have shaped the genome of Corymbia citriodora. Commun Biol. 2021; 4: 1-13

[62]

Myburg AA, Grattapaglia D, Tuskan GA, et al. The genome of Eucalyptus grandis. Nature. 2014; 510: 356-62

[63]

Yuan Z, Fang Y, Zhang T, et al. The pomegranate (Punica granatum L.) genome provides insights into fruit quality and ovule developmental biology . Plant Biotechnol J. 2018; 16: 1363-74

[64]

Gertz EM, Yu Y-K, Agarwala R, et al. Composition-based statistics and translated nucleotide searches: improving the TBLASTN module of BLAST. BMC Biol. 2006; 4: 41

[65]

Wu TD, Watanabe CK . GMAP: a genomic mapping and alignment program for mRNA and EST sequences. Bioinformatics. 2005; 21: 1859-75

[66]

Trapnell C, Pachter L, Salzberg SL . TopHat: discovering splice junctions with RNA-Seq. Bioinformatics. 2009; 25: 1105-11

[67]

Trapnell C, Roberts A, Goff L, et al. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat Protoc. 2012; 7: 562-78

[68]

Haas BJ, Salzberg SL, Zhu W, et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 2008; 9: R7

[69]

Emms DM, Kelly S . OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019; 20: 238

[70]

Fu L, Niu B, Zhu Z, et al. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics. 2012; 28: 3150-2

[71]

Buchfink B, Xie C, Huson DH . Fast and sensitive protein alignment using DIAMOND. Nat Methods. 2015; 12: 59-60

[72]

Wang D, Zhang Y, Zhang Z, et al. KaKs_Calculator 2.0: a toolkit incorporating gamma-series methods and sliding window strategies. Genom Proteom Bioinform. 2010; 8: 77-80

[73]

Zhang Z, Xiao J, Wu J, et al. ParaAT: a parallel tool for constructing multiple protein-coding DNA alignments. Biochem Biophys Res Commun. 2012; 419: 779-81

[74]

Li H . Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018; 34: 3094-100

[75]

Jiang T, Yongzhuang L, Jiang Y, et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 2020; 21: 189

[76]

Marçais G, Delcher AL, Phillippy AM, et al. MUMmer4: a fast and versatile genome alignment system. PLoS Comput Biol. 2018; 14: e1005944

[77]

Dobin A, Davis CA, Schlesinger F, et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013; 29: 15-21

[78]

Liao Y, Smyth GK, Shi W . featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics. 2014; 30: 923- 30

[79]

Love MI, Huber W, Anders S . Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014; 15: 550

[80]

Quinlan AR, Hall IM . BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010; 26: 841-2

PDF (4269KB)

83

Accesses

0

Citation

Detail

Sections
Recommended

/