A chromosome-scale genome assembly for the holly (Ilex polyneura) provides insights into genomic adaptations to elevation in Southwest China

Xin Yao , Zhiqiang Lu , Yu Song , Xiaodi Hu , Richard T. Corlett

Horticulture Research ›› 2022, Vol. 9 ›› Issue (1) : uhab049

PDF (682KB)
Horticulture Research ›› 2022, Vol. 9 ›› Issue (1) :uhab049 DOI: 10.1093/hr/uhab049
Article
research-article
A chromosome-scale genome assembly for the holly (Ilex polyneura) provides insights into genomic adaptations to elevation in Southwest China
Author information +
History +
PDF (682KB)

Abstract

Southwest China is a plant diversity hotspot. The near-cosmopolitan genus Ilex (c. 664 spp., Aquifoliaceae) reaches its maximum diversity in this region, with many narrow-range and a few widespread species. Divergent selection on widespread species leads to local adaptation, with consequences for both conservation and utilization, but is counteracted by geneflow. Many Ilex species are utilized as teas, medicines, ornamentals, honey plants, and timber, but variation below the species level is largely uninvestigated. We therefore studied the widespread Ilex polyneura, which occupies most of the elevational range available and is cultivated for its decorative leafless branches with persistent red fruits. We assembled a chromosome-scale genome using approximately 100x whole genome long-read and short-read sequencing combined with Hi-C sequencing. The genome is approximately 727.1 Mb, with a contig N50 size of 5 124 369 bp and a scaffold N50 size of 36 593 620 bp, for which the BUSCO score was 97.6%, and 98.9% of the assembly was anchored to 20 pseudochromosomes. Out of 32 838 genes predicted, 96.9% were assigned functions. Two whole genome duplication events were identified. Using this genome as a reference, we conducted a population genomics study of 112 individuals from 21 populations across the elevation range using restriction site-associated DNA sequencing (RADseq). Most populations clustered into four clades separated by distance and elevation. Selective sweep analyses identified 34 candidate genes potentially under selection at different elevations, with functions related to responses to abiotic and biotic stresses. This first high-quality genome in the Aquifoliales will facilitate the further domestication of the genus.

Cite this article

Download citation ▾
Xin Yao, Zhiqiang Lu, Yu Song, Xiaodi Hu, Richard T. Corlett. A chromosome-scale genome assembly for the holly (Ilex polyneura) provides insights into genomic adaptations to elevation in Southwest China. Horticulture Research, 2022, 9 (1) : uhab049 DOI:10.1093/hr/uhab049

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Xu X, Dimitrov D, Shrestha N et al. A consistent species richness-climate relationship for oaks across the northern hemisphere. Glob Ecol Biogeogr. 2019; 28: 1051-66.

[2]

Li R, Yue J . A phylogenetic perspective on the evolutionary processes of floristic assemblages within a biodiversity hotspot in eastern Asia. J Syst Evol. 2020; 58: 413-22.

[3]

Yao X, Song Y, Yang JB et al. Phylogeny and biogeography of the hollies (Ilex L., Aquifoliaceae) . J Syst Evol. 2021; 59: 73-82.

[4]

Martins K, Gugger PF, Llanderal-Mendoza J et al. Landscape genomics provides evidence of climate-associated genetic variation in Mexican populations of Quercus rugosa . Evol Appl. 2018; 11: 1842-58.

[5]

Liu Y, Wang H, Jiang Z et al. Genomic basis of geographical adaptation to soil nitrogen in rice. Nature. 2021; 590: 600-5.

[6]

Hong DY . A taxonomical revision of Ilex (Aquifoliaceae) in the pan-Himalaya and unraveling its distribution patterns. Phytotaxa. 2015; 230: 151-71.

[7]

Chen SK, Ma H, Feng Y et al. 2008 Aquifoliaceae. In: Wu ZY, Raven PH, Hong DY, eds. Flora of China.Beijing: Science Press; St. Louis: Missouri Botanical Garden Press, 2008.

[8]

Ming R, Hou S, Feng Y et al. The draft genome of the transgenic tropical fruit tree papaya (Carica papaya Linnaeus) . Nature. 2008; 452: 991-6.

[9]

Luo R, Liu B, Xie Y et al. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience. 2012; 1: 18.

[10]

Hu J, Fan J, Sun Z et al. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics. 2020; 36: 2253-5.

[11]

Roach MJ, Schmidt SA, Borneman AR . Purge Haplotigs: allelic contig reassignment for third-gen diploid genome assemblies. BMC Bioinformatics. 2018; 19: 460.

[12]

Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv 1303.3997. 2013.

[13]

Parra G, Bradnam K, Korf I . CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics. 2007; 23: 1061-7.

[14]

Simão FA, Waterhouse RM, Ioannidis P et al. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 2015; 31: 3210-2.

[15]

Louwers M, Splinter E, Driel RV et al. Studying physical chromatin interactions in plants using chromosome conformation capture (3C). Nat Protoc. 2009; 4: 1216-29.

[16]

Zhang X, Zhang S, Zhao Q et al. Assembly of allele-aware, chromosomal-scale autopolyploid genomes based on Hi-C data. Nat Plants. 2019; 5: 833-45.

[17]

Chen N . Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinformatics. 2004; 5: 4.10.1-4.10.14.

[18]

Zhao X, Wang H . LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 2007; 35: W265-8.

[19]

Price AL, Jones NC, Pevzner PA . De novo identification of repeat families in large genomes . Bioinformatics. 2005; 21: i351-8.

[20]

Benson G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 1999; 27: 573-80.

[21]

Koonin EV, Galperin MY . Sequence - Evolution - Function . Boston, MA: Springer US; 2003.

[22]

Yin J, McLoughlin S, Jeffery IB et al. Integrating multiple genome annotation databases improves the interpretation of microarray gene expression data. BMC Genomics. 2010; 11: 50.

[23]

Mitchell AL, Attwood TK, Babbitt PC et al. InterPro in 2019: improving coverage, classification and access to protein sequence annotations. Nucleic Acids Res. 2019; 47: D351-60.

[24]

Kanehisa M, Sato Y, Kawashima M et al. KEGG as a reference resource for gene and protein annotation. Nucleic Acids Res. 2016; 44: D457-62.

[25]

Stanke M, Diekhans M, Baertsch R et al. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics. 2008; 24: 637-44.

[26]

Blanco E, Parra G, Guigó R . Using geneid to identify genes. Curr Protoc Bioinformatics. 2007; Chapter 4: 4.3.1-4.3.28.

[27]

Burge C, Karlin S . Prediction of complete gene structures in human genomic DNA. J Mol Biol. 1997; 268: 78-94.

[28]

Majoros WH, Pertea M, Salzberg SL . TigrScan and GlimmerHMM: two open-source ab initio eukaryotic gene-finders. Bioinformatics. 2004; 20: 2878-9.

[29]

Li S, Ma L, Li H et al. Snap: an integrated SNP annotation platform. Nucleic Acids Res. 2007; 35: D707-10.

[30]

Hunt SE, Mclaren W, Gil L et al. Ensembl variation resources. Database (Oxford) bay119. 2018.

[31]

Birney E, Durbin R . Using GeneWise in the Drosophila annotation experiment. Genome Res. 2000; 10: 547-8.

[32]

Trapnell C, Pachter L, Salzberg SL . TopHat: discovering splice junctions with RNA-Seq. Bioinformatics. 2009; 25: 1105-11.

[33]

Trapnell C, Roberts A, Goff L et al. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and cufflinks. Nat Protoc. 2012; 7: 562-78.

[34]

Haas BJ, Salzberg SL, Zhu W et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 2008; 9: R7.

[35]

The UniProt Consortium . UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res. 2019; 47: D506-15.

[36]

El-Gebali S, Mistry J, Bateman A et al. The Pfam protein families database in 2019. Nucleic Acids Res. 2019; 47: D427-32.

[37]

Kanehisa M, Furumichi M, Sato Y et al. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 2021; 49: D545-51.

[38]

Hunter S, Apweiler R, Attwood TK et al. InterPro: the integrative protein signature database. Nucleic Acids Res. 2009; 37: D211-5.

[39]

Lowe TM, Chan PP . tRNAscan-SE on-line: integrating search and context for analysis of transfer RNA genes. Nucleic Acids Res. 2016; 44: W54-7.

[40]

Kalvari I, Nawrocki EP, Argasinska J et al. Non-coding RNA analysis using the Rfam database. Curr Protoc Bioinformatics. 2018; 62: e51.

[41]

Wang Y, Tang H, Debarry JD et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res. 2012; 40: e49.

[42]

Edgar RC . MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004; 32: 1792-7.

[43]

Yang Z . PAML 4: phylogenetic analysis by maximum likelihood. Mol Biol Evol. 2007; 24: 1586-91.

[44]

Baird NA, Etter PD, Atwood TS et al. Rapid SNP discovery and genetic mapping using sequenced RAD markers. PLoS One. 2008; 3: e3376.

[45]

Bolger AM, Lohse M, Usadel B . Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014; 30: 2114-20.

[46]

Li H, Handsaker B, Wysoker A et al. The sequence alignment/map format and SAMtools. Bioinformatics. 2009; 25: 2078-9.

[47]

DePristo M, Banks E, Poplin R et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nat Genet. 2011; 43: 491-8.

[48]

Yang J, Lee SH, Goddard ME et al. GCTA: a tool for genome-wide complex trait analysis. Am J Hum Genet. 2011; 88: 76-82.

[49]

Purcell S, Neale B, Todd-Brown K et al. PLINK: a toolset for whole-genome association and population-based linkage analysis. Am J Hum Genet. 2007; 81: 559-75.

[50]

Tang H, Peng J, Wang P et al. Estimation of individual admixture: analytical and study design considerations. Genet Epidemiol. 2005; 28: 289-301.

[51]

Zhang C, Dong SS, Xu JJ et al. PopLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics. 2019; 35: 1786-8.

[52]

Danecek P, Auton A, Abecasis G et al. The variant call format and VCFtools. Bioinformatics. 2011; 27: 2156-8.

[53]

Pu X, Li Z, Tian Y et al. The honeysuckle genome provides insight into the molecular mechanism of carotenoid metabolism underlying dynamic flower coloration. New Phytol. 2020; 227: 930-43.

[54]

Song X, Wang J, Li N et al. Deciphering the high-quality genome sequence of coriander that causes controversial feelings. Plant Biotechnol J. 2020; 18: 1444-56.

[55]

He S, Dong X, Zhang G et al. High quality genome of Erigeron breviscapus provides a reference for herbal plants in Asteraceae. Mol Ecol Resour. 2021; 21: 153-69.

[56]

Fan DM, Yue JP, Nie ZL et al. Phylogeography of Sophora davidii (Leguminosae) across the ‘Tanaka-Kaiyong line’, an important phytogeographic boundary in Southwest China. Mol Ecol. 2013; 22: 4270-88.

[57]

Qian LS, Chen JH, Deng T et al. Plant diversity in Yunnan: current status and future directions. Plant Divers. 2020; 42: 281-91.

[58]

Chen J, Huang Y, Brachi B et al. Genome-wide analysis of cushion willow provides insights into alpine plant divergence in a biodiversity hotspot. Nat Commun. 2019; 10: 5230.

[59]

Hasanuzzaman M, Nahar K, Anee TI et al. Glutathione in plants: biosynthesis and physiological role in environmental stress tolerance. Physiol Mol Biol Plants. 2017; 23: 249-68.

[60]

Vriese KD, Pollier J, Goossens A et al. Dissecting cholesterol and phytosterol biosynthesis via mutants and inhibitors. J Exp Bot. 2021; 72: 241-53.

[61]

Kang K, Yue L, Xia X et al. Comparative metabolomics analysis of different resistant rice varieties in response to the brown planthopper Nilaparvata lugens Hemiptera: Delphacidae. Metabolomics. 2019; 15: 62.

[62]

Zhang L, Paasch BC, Chen J et al. An important role of l-fucose biosynthesis and protein fucosylation genes in Arabidopsis immunity. New Phytol. 2019; 222: 981-94.

PDF (682KB)

57

Accesses

0

Citation

Detail

Sections
Recommended

/