The chromosome-scale assembly of the willow genome provides insight into Salicaceae genome evolution

Suyun Wei , Yonghua Yang , Tongming Yin

Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) : 45

PDF (3062KB)
Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) :45 DOI: 10.1038/s41438-020-0268-6
Article
research-article
The chromosome-scale assembly of the willow genome provides insight into Salicaceae genome evolution
Author information +
History +
PDF (3062KB)

Abstract

Salix suchowensis is an early-flowering shrub willow that provides a desirable system for studies on the basic biology of woody plants. The current reference genome of S. suchowensis was assembled with 454 sequencing reads. Here, we report a chromosome-scale assembly of S. suchowensis generated by combining PacBio sequencing with Hi-C technologies. The obtained genome assemblies covered a total length of 356 Mb. The contig N50 of these assemblies was 263,908 bp, which was ~65-fold higher than that reported previously. The contiguity and completeness of the genome were significantly improved. By applying Hi-C data, 339.67 Mb (95.29%) of the assembled sequences were allocated to the 19 chromosomes of haploid willow. With the chromosome-scale assembly, we revealed a series of major chromosomal fissions and fusions that explain the genome divergence between the sister genera of Salix and Populus. The more complete and accurate willow reference genome obtained in this study provides a fundamental resource for studying many genetic and genomic characteristics of woody plants.

Cite this article

Download citation ▾
Suyun Wei, Yonghua Yang, Tongming Yin. The chromosome-scale assembly of the willow genome provides insight into Salicaceae genome evolution. Horticulture Research, 2020, 7 (1) : 45 DOI:10.1038/s41438-020-0268-6

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Tuskan, G. A. et al. The genome of black cottonwood, Populus trichocarpa (Torr. & Gray). Science 313, 1596-1604 (2006).

[2]

Ma, T. et al. Genomic insights into salt adaptation in a desert poplar. Nat. Commun. 4, 2797 (2013).

[3]

Yang, W. et al. The draft genome sequence of a desert tree Populus pruinosa. GigaScience 6, gix075 (2017).

[4]

Ma, J. et al. Genome sequence and genetic transformation of a widely distributed and cultivated poplar. Plant Biotechnol. J. 17, 451-460 (2019).

[5]

Liu, Y., Wang, X. & Zeng, Q. De novo assembly of white poplar genome and genetic diversity of white poplar population in Irtysh River basin in China. Sci. China Life Sci. 62, 609-618 (2019).

[6]

Lin, Y. et al. Functional and evolutionary genomic inferences in Populus through genome and population sequencing of American and European aspen. Proc. Natl Acad. Sci. 115, E10970-E10978 (2018).

[7]

Nordberg, H. et al. The genome portal of the Department of Energy Joint Genome Institute: 2014 updates. Nucleic Acids Res. 42, D26-D31 (2013).

[8]

Dai, X. et al. The willow genome and divergent evolution from poplar after the common genome duplication. Cell Res. 24, 1274 (2014).

[9]

Eid, J. et al. Real-time DNA sequencing from single polymerase molecules. Science 323, 133-138 (2009).

[10]

Dekker, J . The three’C’s of chromosome conformation capture: controls, controls, controls. Nat. Methods 3, 17 (2006).

[11]

Zhang, L. et al. Improved Brassica rapa reference genome by single-molecule sequencing and chromosome conformation capture technologies. Horticulture Res. 5, 50 (2018).

[12]

Jiao, Y. et al. Improved maize reference genome with single-molecule technologies. Nature 546, 524 (2017).

[13]

Jibran, R. et al. Chromosome-scale scaffolding of the black raspberry (Rubus occidentalis L.) genome based on chromatin interaction data. Horticulture Res. 5, 8 (2018).

[14]

Echenwalder, J. (eds). Systematics and Evolution of Populus. Biology of Populus Ands It Implications for Management and Conservation. Part I. (NRC Research Press, 1996).

[15]

Argus, G. W. Infrageneric classification of Salix (Salicaceae) in the new world. Syst. Bot. Monogr. 52, 1-121 (1997).

[16]

Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210-3212 (2015).

[17]

Kohler, A. et al. Genome-wide identification of NBS resistance genes in Populus trichocarpa. Plant Mol. Biol. 66, 619-636 (2008).

[18]

Bresson, A. et al. Qualitative and quantitative resistances to leaf rust finely mapped within two nucleotide-binding site leucine-rich repeat (NBS-LRR)-rich genomic regions of chromosome 19 in poplar. N. Phytologist 192, 151-163 (2011).

[19]

Santner, A. & Estelle, M. Recent advances and emerging trends in plant hormone signalling. Nature 459, 1071 (2009).

[20]

Hou, J., Wei, S., Pan, H., Zhuge, Q. & Yin, T. Uneven selection pressure accelerating divergence of Populus and Salix. Horticulture Res. 6, 37 (2019).

[21]

Hou, J. et al. Major chromosomal rearrangements distinguish willow and poplar after the ancestral “Salicoid” genome duplication. Genome Biol. Evol. 8, 1868-1875 (2016).

[22]

Hanley, S., Mallott, M. & Karp, A. Alignment of a Salix linkage map to the Populus genomic sequence reveals macrosynteny between willow and poplar genomes. Tree Genet. Genomes 3, 35-48 (2006).

[23]

Berlin, S., Lagercrantz, U., von Arnold, S., Öst, T. & Rönnberg-Wästljung, A. C. High-density linkage mapping and evolution of paralogs and orthologs in Salix and Populus. BMC Genomics 11, 129 (2010).

[24]

Hou, J. et al. Different autosomes evolved into sex chromosomes in the sister genera of Salix and Populus. Sci. Rep. 5, 9076 (2015).

[25]

Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722-736 (2017).

[26]

Chaisson, M. J. & Tesler, G. Mapping single molecule sequencing reads using basic local alignment with successive refinement (BLASR): application and theory. BMC Bioinforma. 13, 238 (2012).

[27]

Li, H. Toward better understanding of artifacts in variant calling from high-coverage samples. Bioinformatics 30, 2843-2851 (2014).

[28]

Li, H. et al. The sequence alignment/map format and SAMtools. Bioinformatics 25, 2078-2079 (2009).

[29]

Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16, 259 (2015).

[30]

Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760 (2009).

[31]

Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119 (2013).

[32]

Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, W265-W268 (2007).

[33]

Price, A. L., Jones, N. C. & Pevzner, P. A. De novo identification of repeat families in large genomes. Bioinformatics 21, i351-i358 (2005).

[34]

Edgar, R. C. & Myers, E. W. PILER: identification and classification of genomic repeats. Bioinformatics 21, i152-i158 (2005).

[35]

Wicker, T. et al. A unified classification system for eukaryotic transposable elements. Nat. Rev. Genet. 8, 973 (2007).

[36]

Tarailo-Graovac, M. & Chen, N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinforma. 25, 4.10.11-14.10.14 (2009).

[37]

Jurka, J. et al. Repbase Update, a database of eukaryotic repetitive elements. Cytogenetic Genome Res. 110, 462-467 (2005).

[38]

Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J. Basic local alignment search tool. J. Mol. Biol. 215, 403-410 (1990).

[39]

Keilwagen, J. et al. Using intron position conservation for homology-based gene prediction. Nucleic Acids Res. 44, e89- e89 (2016).

[40]

Kim, D., Langmead, B. & Salzberg, S. L. HISAT: a fast spliced aligner with low memory requirements. Nat. Methods 12, 357 (2015).

[41]

Tang, S., Lomsadze, A. & Borodovsky, M. Identification of protein coding regions in RNA transcripts. Nucleic Acids Res. 43, e78- e78 (2015).

[42]

Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the Program to Assemble Spliced Alignments. Genome Biol. 9, R7 (2008).

[43]

Campbell, M. A., Haas, B. J., Hamilton, J. P., Mount, S. M. & Buell, C. R. Comprehensive analysis of alternative splicing in rice and comparative analyses with Arabidopsis. BMC Genomics 7, 327 (2006).

[44]

Camacho, C. et al. BLAST+: architecture and applications. BMC Bioinforma. 10, 421 (2009).

[45]

Marchler-Bauer, A. et al. CDD: a Conserved Domain Database for the functional annotation of proteins. Nucleic Acids Res. 39, D225-D229 (2010).

[46]

Dimmer, E. C. et al. The UniProt-GO annotation database in 2011. Nucleic Acids Res. 40, D565-D570 (2011).

[47]

Mitchell, A. et al. The InterPro protein families database: the classification resource after 15 years. Nucleic Acids Res. 43, D213-D221 (2014).

[48]

Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res. 25, 955-964 (1997).

[49]

Nawrocki, E. P. & Eddy, S. R. Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics 29, 2933-2935 (2013).

[50]

Griffiths-Jones, S., Grocock, R. J., Van Dongen, S., Bateman, A. & Enright, A. J. miRBase: microRNA sequences, targets and gene nomenclature. Nucleic Acids Res. 34, D140-D144 (2006).

[51]

Griffiths-Jones, S. et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 33, D121-D124 (2005).

[52]

Marçais, G. et al. MUMmer4: a fast and versatile genome alignment system. PLoS Comput. Biol. 14, e1005944 (2018).

[53]

Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 (2013).

[54]

Emms, D. M. & Kelly, S. OrthoFinder: solving fundamental biases in whole genome comparisons dramatically improves orthogroup inference accuracy. Genome Biol. 16, 157 (2015).

[55]

Yu, G., Wang, L. G., Han, Y. & He, Q. Y. clusterProfiler: an R package for comparing biological themes among gene clusters. Omics: J. Integr. Biol. 16, 284-287 (2012).

[56]

Finn, R. D., Clements, J. & Eddy, S. R. HMMER web server: interactive sequence similarity searching. Nucleic Acids Res. 39, W29-W37 (2011).

[57]

Letunic, I. & Bork, P. 20 years of the SMART protein domain annotation resource. Nucleic Acids Res. 46, D493-D496 (2017).

[58]

Larkin, M. A. et al. Clustal W and clustal X version 2.0. Bioinformatics 23, 2947-2948 (2007).

[59]

Kumar, S., Stecher, G., Li, M., Knyaz, C. & Tamura, K. MEGA X: molecular evolutionary genetics analysis across computing platforms. Mol. Biol. Evol. 35, 1547-1549 (2018).

[60]

Letunic, I. & Bork, P. Interactive tree of life (iTOL) v3: an online tool for the display and annotation of phylogenetic and other trees. Nucleic Acids Res. 44, W242-W245 (2016).

[61]

Wang, D., Zhang, Y., Zhang, Z., Zhu, J. & Yu, J. KaKs_Calculator 2.0: a toolkit incorporating gamma-series methods and sliding window strategies. Genomics, Proteom. Bioinforma. 8, 77-80 (2010).

[62]

Krzywinski, M. et al. Circos: an information aesthetic for comparative genomics. Genome Res. 19, 1639-1645 (2009).

PDF (3062KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/