A high-quality reference genome of wild Cannabis sativa

Shan Gao , Baishi Wang , Shanshan Xie , Xiaoyu Xu , Jin Zhang , Li Pei , Yongyi Yu , Weifei Yang , Ying Zhang

Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) : 73

PDF (948KB)
Horticulture Research ›› 2020, Vol. 7 ›› Issue (1) :73 DOI: 10.1038/s41438-020-0295-3
Article
research-article
A high-quality reference genome of wild Cannabis sativa
Author information +
History +
PDF (948KB)

Abstract

Cannabis sativa is a well-known plant species that has great economic and ecological significance. An incomplete genome of cloned C. sativa was obtained by using SOAPdenovo software in 2011. To further explore the utilization of this plant resource, we generated an updated draft genome sequence for wild-type varieties of C. sativa in China using PacBio single-molecule sequencing and Hi-C technology. Our assembled genome is approximately 808 Mb, with scaffold and contig N50 sizes of 83.00 Mb and 513.57 kb, respectively. Repetitive elements account for 74.75% of the genome. A total of 38,828 protein-coding genes were annotated, 98.20% of which were functionally annotated. We provide the first comprehensive de novo genome of wild-type varieties of C. sativa distributed in Tibet, China. Due to long-term growth in the wild environment, these varieties exhibit higher heterozygosity and contain more genetic information. This genetic resource is of great value for future investigations of cannabinoid metabolic pathways and will aid in promoting the commercial production of C. sativa and the effective utilization of cannabinoids. The assembled genome is also a valuable resource for intensively and effectively investigating the C. sativa genome further in the future.

Cite this article

Download citation ▾
Shan Gao, Baishi Wang, Shanshan Xie, Xiaoyu Xu, Jin Zhang, Li Pei, Yongyi Yu, Weifei Yang, Ying Zhang. A high-quality reference genome of wild Cannabis sativa. Horticulture Research, 2020, 7 (1) : 73 DOI:10.1038/s41438-020-0295-3

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Schultes, R. E., Klein, W. M., Plowman, T. & Lockwood, T. E. Cannabis: an example of taxonomic neglect. Bot. Mus. Leafl., Harv. Univ. 23, 337-367 (1974).

[2]

Li, H.-L. An archaeological and historical account of cannabis in China. Econ. Bot. 28, 437-448 (1973).

[3]

Leung, L. Cannabis and its derivatives: review of medical use. J. Am. Board Fam. Med. 24, 452-462 (2011).

[4]

Ruiz, L., Miguel, A. & Díaz-Laviada, I. Δ9‐Tetrahydrocannabinol induces apoptosis in human prostate PC‐3 cells via a receptor‐independent mechanism. FEBS Lett. 458, 400-404 (1999).

[5]

Esposito, G., De Filippis, D., Carnuccio, R., Izzo, A. A. & Iuvone, T. The marijuana component cannabidiol inhibits β-amyloid-induced tau protein hyperphosphorylation through Wnt/β-catenin pathway rescue in PC12 cells. J. Mol. Med. 84, 253-258 (2006).

[6]

Martín-Moreno, A. M. et al. Cannabidiol and other cannabinoids reduce microglial activation in vitro and in vivo: relevance to Alzheimers′ disease. Mol. Pharmacol. 79, 964-973 (2011).

[7]

Steffens, S. et al. Low dose oral cannabinoid therapy reduces progression of atherosclerosis in mice. Nature 434, 782 (2005).

[8]

Taura, F., Sirikantaramas, S., Shoyama, Y., Shoyama, Y. & Morimoto, S. Phytocannabinoids in Cannabis sativa: recent studies on biosynthetic enzymes. Chem. Biodivers. 4, 1649-1663 (2007).

[9]

Van Bakel, H. et al. The draft genome and transcriptome of Cannabis sativa. Genome Biol. 12, R102 (2011).

[10]

Ming, R., Bendahmane, A. & Renner, S. S. Sex chromosomes in land plants. Annu. Rev. Plant Biol. 62, 485-514 (2011).

[11]

Sakamoto, K., Akiyama, Y., Fukui, K., Kamada, H. & Satoh, S. Characterization; genome sizes and morphology of sex chromosomes in hemp (Cannabis sativa L.). Cytologia 63, 459-464 (1998).

[12]

Li, Y. H. et al. De novo assembly of soybean wild relatives for pan-genome analysis of diversity and agronomic traits. Nat. Biotechnol. 32, 1045-1052 (2014).

[13]

Walker, B. J. et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One 9, e112963 (2014).

[14]

Lieberman-Aiden, E. et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science 326, 289-293 (2009).

[15]

Rao, S. S. et al. A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell 159, 1665-1680 (2014).

[16]

Belaghzal, H., Dekker, J. & Gibcus, J. H. Hi-C 2.0: An optimized Hi-C procedure for high-resolution genome-wide mapping of chromosome conformation. Methods 123, 56-65 (2017).

[17]

Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119-1125 (2013).

[18]

Simao, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210-3212 (2015).

[19]

Kong, L. et al. CPC: assess the protein-coding potential of transcripts using sequence features and support vector machine. Nucleic Acids Res. 35, W345-W349 (2007).

[20]

Finn, R. D. et al. Pfam: the protein families database. Nucleic Acids Res. 42, D222-D230 (2014).

[21]

Powell, S. et al. eggNOG v3.0: orthologous groups covering 1133 organisms at 41 different taxonomic ranges. Nucleic Acids Res. 40, D284-D289 (2012).

[22]

Ashburner, M. et al. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nat. Genet. 25, 25-29 (2000).

[23]

Kanehisa, M. Molecular network analysis of diseases and drugs in KEGG. Methods Mol. Biol. 939, 263-275 (2013).

[24]

Griffiths-Jones, S. et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 33, D121-D124 (2005).

[25]

Lowe, T. M. & Eddy, S. R. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res. 25, 955-964 (1997).

[26]

Li, L., Stoeckert, C. J. Jr. & Roos, D. S. OrthoMCL: identification of ortholog groups for eukaryotic genomes. Genome Res. 13, 2178-2189 (2003).

[27]

Edgar, R. C. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 32, 1792-1797 (2004).

[28]

Guindon, S. et al. New algorithms and methods to estimate maximum-likelihood phylogenies: assessing the performance of PhyML 3.0. Syst. Biol. 59, 307-321 (2010).

[29]

Yang, Z. PAML 4: phylogenetic analysis by maximum likelihood. Mol. Biol. Evol. 24, 1586-1591 (2007).

[30]

Kumar, S., Stecher, G., Suleski, M. & Hedges, S. B. TimeTree: a resource for timelines, timetrees, and divergence times. Mol. Biol. Evol. 34, 1812-1819 (2017).

[31]

Zeng, Q. et al. Definition of eight mulberry species in the genus morus by internal transcribed spacer-based phylogeny. PLoS One 10, e0135411 (2015).

[32]

Foster, C. S. P. et al. Evaluating the impact of genomic data and priors on Bayesian estimates of the angiosperm evolutionary timescale. Syst. Biol. 66, 338-351 (2017).

[33]

Massoni, J., Couvreur, T. L. & Sauquet, H. Five major shifts of diversification through the long evolutionary history of Magnoliidae (angiosperms). BMC Evolut. Biol. 15, 49 (2015).

[34]

De Bie, T., Cristianini, N., Demuth, J. P. & Hahn, M. W. CAFE: a computational tool for the study of gene family evolution. Bioinformatics 22, 1269-1271 (2006).

[35]

Huang, S. et al. The genome of the cucumber, Cucumis sativus L. Nat. Genet. 41, 1275-1281 (2009).

[36]

Schmutz, J. et al. Genome sequence of the palaeopolyploid soybean. Nature 463, 178-183 (2010).

[37]

Jaillon, O. et al. The grapevine genome sequence suggests ancestral hexaploidization in major angiosperm phyla. Nature 449, 463-467, (2007).

[38]

Tang, H. et al. Synteny and collinearity in plant genomes. Science 320, 486-488 (2008).

[39]

Soderlund, C., Bomhoff, M. & Nelson, W. M. SyMAP v3.4: a turnkey synteny system with application to plant genomes. Nucleic Acids Res. 39, e68 (2011).

[40]

Dujon, B. et al. Genome evolution in yeasts. Nature 430, 35-44 (2004).

[41]

Shi, J. et al. Chromosome conformation capture resolved near complete genome assembly of broomcorn millet. Nature 10, 464 (2019).

[42]

Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinform. 9, 18 (2008).

[43]

Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, W265-W268 (2007).

[44]

OuS. & JiangN. LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410-1422 (2018).

[45]

English, A. C. et al. Mind the gap: upgrading genomes with Pacific Biosciences RS long-read sequencing technology. PLoS One 7, e47768 (2012).

[46]

Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 27, 722-736 (2017).

[47]

Istace, B. et al. De novo assembly and population genomic survey of natural yeast isolates with the Oxford Nanopore MinION sequencer. GigaScience 6, 1-13 (2017).

[48]

Chakraborty, M., Baldwin-Brown, J. G., Long, A. D. & Emerson, J. J. Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage. Nucleic Acids Res. 44, e147 (2016).

[49]

Servant, N. et al. HiTC: exploration of high-throughput ‘C’ experiments. Bioinformatics 28, 2843-2844 (2012).

[50]

Jurka, J. et al. Repbase Update, a database of eukaryotic repetitive elements. Cytogenet. Genome Res. 110, 462-467 (2005).

[51]

Keilwagen, J. et al. Using intron position conservation for homology-based gene prediction. Nucleic Acids Res. 44, e89 (2016).

PDF (948KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/