A chromosome-level genome assembly of Agave hybrid NO.11648 provides insights into the CAM photosynthesis

Ziping Yang , Qian Yang , Qi Liu , Xiaolong Li , Luli Wang , Yanmei Zhang , Zhi Ke , Zhiwei Lu , Huibang Shen , Junfeng Li , Wenzhao Zhou

Horticulture Research ›› 2024, Vol. 11 ›› Issue (2) : 269

PDF (2195KB)
Horticulture Research ›› 2024, Vol. 11 ›› Issue (2) :269 DOI: 10.1093/hr/uhad269
Articles
research-article
A chromosome-level genome assembly of Agave hybrid NO.11648 provides insights into the CAM photosynthesis
Author information +
History +
PDF (2195KB)

Abstract

The subfamily Agavoideae comprises crassulacean acid metabolism (CAM), C3, and C4 plants with a young age of speciation and slower mutation accumulation, making it a model crop for studying CAM evolution. However, the genetic mechanism underlying CAM evolution remains unclear because of lacking genomic information. This study assembled the genome of Agave hybrid NO.11648, a constitutive CAM plant belonging to subfamily Agavoideae, at the chromosome level using data generated from high-throughput chromosome conformation capture, Nanopore, and Illumina techniques, resulting in 30 pseudo-chromosomes with a size of 4.87 Gb and scaffold N50 of 186.42 Mb. The genome annotation revealed 58 841 protein-coding genes and 76.91% repetitive sequences, with the dominant repetitive sequences being the I-type repeats (Copia and Gypsy accounting for 18.34% and 13.5% of the genome, respectively). Our findings also provide support for a whole genome duplication event in the lineage leading to A. hybrid, which occurred after its divergence from subfamily Asparagoideae. Moreover, we identified a gene duplication event in the phosphoenolpyruvate carboxylase kinase (PEPCK) gene family and revealed that three PEPCK genes (PEPCK3, PEPCK5, and PEPCK12) were involved in the CAM pathway. More importantly, we identified transcription factors enriched in the circadian rhythm, MAPK signaling, and plant hormone signal pathway that regulate the PEPCK3 expression by analysing the transcriptome and using yeast one-hybrid assays. Our results shed light on CAM evolution and offer an essential resource for the molecular breeding program of Agave spp.

Cite this article

Download citation ▾
Ziping Yang, Qian Yang, Qi Liu, Xiaolong Li, Luli Wang, Yanmei Zhang, Zhi Ke, Zhiwei Lu, Huibang Shen, Junfeng Li, Wenzhao Zhou. A chromosome-level genome assembly of Agave hybrid NO.11648 provides insights into the CAM photosynthesis. Horticulture Research, 2024, 11 (2) : 269 DOI:10.1093/hr/uhad269

登录浏览全文

4963

注册一个新账户 忘记密码

Acknowledgements

This study was sponsored by the Earmarked fund for the China Agriculture Research System (grant No. CARS-19), the National Natural Science Foundation of China (grant No. 31801679), Guangdong Provincial Team of Technical System Innovation for Sugarcane Sisal Hemp Industry (grant No. 2023KJ104-03), Guangdong Basic and Applied Basic Research Foundation (grant Nos 2021A1515012421 and 2022A1515011841), Hainan Provincial Natural Science Foundation of China (321QN300 and 323MS099), and Central Public-interest Scientific Institution Basal Research Fund for Chinese Academy of Tropical Agricultural Sciences (grant Nos. 1630062019016, 1630062020015, 1630062022002, and 1630062021015).

Author contributions

Z.Y. and W.Z. designed and coordinated the whole project. Z.Y. and Q.Y. led and performed the whole project. Q.L., X.L., L.W., and Y.Z. analysed the data. Z.K., Z.L., and H.S. collected and processed the samples. Z.Y. and Q.Y. drafted the manuscript. W.Z. revised the manuscript. Z.K., Z.L., H.S., and J.L. contributed to the manuscript preparation. All authors read and approved the final manuscript.

Data availability

Data generated during this study have been deposited into National Genomics Data Center under the project accession number PRJCA016359 (https://ngdc.cncb.ac.cn/gsub/).

Conflict of interest statement

The authors declare that they have no conflict of interest.

Supplementary data

Supplementary data is available at Horticulture Research online.

References

[1]

Good-Avila SV, Souza V, Gaut BS. et al. Timing and rate of speciation in Agave.(Agavaceae). Proc Natl Acad Sci U S A. 2006; 103: 9124-9

[2]

The Angiosperm Phylogeny Group and others. An update of the Angiosperm Phylogeny Group classification for the orders and families of flowering plants: APGIV. Bot J Linn Soc. 2016; 181:1-20

[3]

Davis SC, Simpson J, Gil-Vega KC. et al. Undervalued potential of crassulacean acid metabolism for current and future agricultural production. J Exp Bot. 2019; 70:6521-37

[4]

Stewart JR. Agave as a model CAM crop system for a warming and drying world. Front Plant Sci. 2015; 6:684

[5]

Trejo L, Limones V, Peña G. et al. Genetic variation and relationships among agaves related to the production of tequila and mezcal in Jalisco. Ind Crop Prod. 2018; 125:140-9

[6]

Yang XH, Cushman JC, Borland AM. et al. A roadmap for research on crassulacean acid metabolism (CAM) to enhance sustainable food and bioenergy production in a hotter, drier world. New Phytol. 2015; 207:491-504

[7]

Silvera K, Neubig KM, Whitten WM. et al. Evolution along the crassulacean acid metabolism continuum. Funct Plant Biol. 2010; 37:995-1010

[8]

Wickell D, Kuo LY, Yang HP. et al. Underwater CAM photosynthesis elucidated by Isoetes genome. Nat Commun. 2021; 12:6348

[9]

Cai J, Liu X, Vanneste K. et al. The genome sequence of the orchid Phalaenopsis equestris. Nat Genet. 2015; 47:65-72

[10]

West-Eberhard MJ, Smith JAC, Winter K. Photosynthesis, reorganized. Science. 2011; 332:311-2

[11]

Heyduk K, Ray JN, Ayyampalayam S. et al. Shifts in gene expression profiles are associated with weak and strong Crassulacean acid metabolism. Am J Bot. 2018; 105:587-601

[12]

Ming R, VanBuren R, Wai CM. et al. The pineapple genome and the evolution of CAM photosynthesis. Nat Genet. 2015; 47:1435-42

[13]

Yang XH, Hu R, Yin H. et al. The Kalanchoë genome provides insights into convergent evolution and building blocks of crassulacean acid metabolism. Nat Commun. 2017; 8:1899

[14]

Yin HF, Guo HB, Weston DJ. et al. Diel rewiring and positive selection of ancient plant proteins enabled evolution of CAM photosynthesis in Agave. BMC Genomics. 2018; 19:588

[15]

Robert ML, Lim KY, Hanson L. et al. Wild and agronomically important Agave species (Asparagaceae) show proportional increases in chromosome number, genome size, and genetic markers with increasing ploidy. Bot J Lin Soc. 2008; 158:215-22

[16]

Bousios A, Saldana-Oyarzabal I, Valenzuela-Zapata AG. et al. Isolation and characterization of Ty1-copia retrotransposon sequences in the blue agave (Agave tequilana Weber var. Azul) and their development as SSAP markers for phylogenetic analysis. Plant Sci. 2007; 172:291-8

[17]

Sandoval S d CD, Juárez MJA, Simpson J. Agave tequilana MADS genes show novel expression patterns in meristems, developing bulbils and floral organs. Sex Plant Reprod. 2012; 25:11-26

[18]

Sun XD, Zhu S, Li N. et al. A chromosome-level genome assembly of garlic (Allium sativum) provides insights into genome evolution and allicin biosynthesis. Mol Plant. 2020; 13:1328-39

[19]

Cheng H, Song X, Hu Y. et al. Chromosome-level wild Hevea brasiliensis genome provides new tools for genomic-assisted breeding and valuable loci to elevate rubber yield. Plant Biotechnol J. 2023; 21:1058-72

[20]

Castorena-Sánchez I, Escobedo RM, Quiroz A. New cytotaxonomical determinants recognized in six taxa of Agave in the sections Rigidae and Sisalanae. Can J Bot. 1991; 69:1257-64

[21]

Ou CQ, Wang F, Wang J. et al. A de novo genome assembly of the dwarfing pear rootstock Zhongai 1. Sci Data. 2019; 6:281

[22]

Wu HL, Ma T, Kang M. et al. A high-quality Actinidia chinensis (kiwifruit) genome. Hortic Res. 2019; 6:117

[23]

Trapnell C, Pachter L, Salzberg SL. TopHat: discovering splice junctions with RNA-Seq. Bioinformatics. 2009; 25:1105-11

[24]

Deng H, Zhang LS, Zhang GQ. et al. Evolutionary history of PEPC genes in green plants: implications for the evolution of CAM in orchids. Mol Phylogenet Evol. 2016; 94:559-64

[25]

Zhang LS, Chen F, Zhang GQ. et al. Origin and mechanism of crassulacean acid metabolism in orchids as implied by comparative transcriptomics and genomics of the carbon fixation pathway. Plant J. 2016; 86:175-85

[26]

Heyduk K, Mckain MR, Lalani F. et al. Evolution of a CAM anatomy predates the origins of Crassulacean acid metabolism in the Agavoideae (Asparagaceae). Mol Phylogenet Evol. 2016; 105:102-13

[27]

Liu B, Shi Y, Yuan J. et al. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. arXiv. 2013. preprint: not peer reviewed

[28]

Koren S, Walenz BP, Berlin K. et al. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res. 2017; 27:722-36

[29]

Ruan J, Li H. Fast and accurate long-read assembly with Wtdbg2. Nat Methods. 2019; 17:155-8

[30]

Vaser R, Sovic I, Nagarajan N. et al. Fast and accurate de novo genome assembly from long uncorrected reads. Genome Res. 2017; 27:737-46

[31]

Walker BJ, Abeel T, Shea T. et al. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PLoS One. 2014; 9:e112963

[32]

Roach MJ, Schmidt SA, Borneman AR. Purge Haplotigs: allelic contig reassignment for third-gen diploid genome assemblies. BMC Bioinformatics. 2018; 19:460

[33]

Chakraborty M, Baldwin-Brown JG, Long AD. et al. Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage. Nucl Acids Res. 2016; 44:e147

[34]

Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv. 2013. preprint: not peer reviewed

[35]

Simão FA, Waterhouse RM, Ioannidis P. et al. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics. 2015; 31:3210-2

[36]

Ou SJ, Chen JF, Jiang N. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic Acids Res. 2018; 46:e126

[37]

Li H, Durbin R. Fast and accurate short read alignment with burrows-wheeler transform. Bioinformatics. 2009; 25:1754-60

[38]

Servant N, Varoquaux N, Lajoie BR. et al. HiC-pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 2015; 16:259

[39]

Burton JN, Adey A, Patwardhan RP. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat Biotechnol. 2013; 31:1119-25

[40]

Pertea M, Pertea GM, Antonescu CM. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat Biotechnol. 2015; 33:290-5

[41]

Kim D, Langmead B, Salzberg SL. HISAT: a fast spliced aligner with low memory requirements. Nat Methods. 2015; 12:357-60

[42]

Grabherr MG, Haas BJ, Yassour M. et al. Trinity: reconstructing a full-length transcriptome without a genome from RNA-Seq data. Nat Biotechnol. 2011; 29:644-52

[43]

Campbell MA, Haas BJ, Hamilton JP. et al. Comprehensive analysis of alternative splicing in rice and comparative analyses with Arabidopsis. BMC Genomics. 2006; 7:327

[44]

Li SF, Wang J, Dong R. et al. Chromosome-level genome assembly, annotation and evolutionary analysis of the ornamental plant Asparagus setaceus. Hortic Res. 2020; 7:48

[45]

Harkess A, Zhou J, Xu C. et al. The asparagus genome sheds light on the origin and evolution of a young Y chromosome. Nat Commun. 2017; 8:1279

[46]

Ouyang S, Zhu W, Hamilton J. et al. The TIGR rice genome annotation resource: improvements and new features. Nucleic Acids Res. 2006; 35:D883-7

[47]

Altschul SF, Madden TL, Schäffer AA. et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997; 25:3389-402

[48]

Keilwagen J, Wenk M, Erickson JL. et al. Using intron position conservation for homology-based gene prediction. Nucleic Acids Res. 2016; 44:e89

[49]

Keilwagen J, Hartung F, Paulini M. et al. Combining RNA-seq data and homology-based gene prediction for plants, animals and fungi. BMC Bioinformatics. 2018; 19:189

[50]

Burge C, Karlin S. Prediction of complete gene structures in human genomic DNA. J Mol Biol. 1997; 268:78-94

[51]

Alioto T, Blanco E, Parra G. et al. Using geneid to identify Genes. Curr Protoc Bioinformatics. 2018; 64:e56

[52]

Stanke M, Waack S. Gene prediction with a hidden Markov model and a new intron submodel. Bioinformatics. 2003; 19:ii215-225

[53]

Majoros WH, Pertea M, Salzberg SL. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics. 2004; 20:2878-9

[54]

Korf I. Gene finding in novel genomes. BMC Bioinformatics. 2004; 5:59

[55]

Haas BJ, Salzberg SL, Zhu W. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 2008; 9:R7

[56]

Kanehisa M, Goto S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 2000; 28:27-30

[57]

Koonin EV, Fedorova ND, Jackson JD. et al. A comprehensive evolutionary classification of proteins encoded in complete eukaryotic genomes. Genome Biol. 2004; 5:R7

[58]

Boeckmann B, Bairoch A, Apweiler R. et al. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003. Nucleic Acids Res. 2003; 31:365-70

[59]

Marchler-Bauer A, Lu S, Anderson JB. et al. CDD: a Conserved Domain Database for the functional annotation of proteins. Nucleic Acids Res. 2011; 39:D225-9

[60]

Dimmer EC, Huntley RP, Alam-Faruque Y. et al. The UniProt-GO annotation database in 2011. Nucleic Acids Res. 2012; 40:D565-70

[61]

Altschul SF, Gish W, Miller W. et al. Basic local alignment search tool. J Mol Biol. 1990; 215:403-10

[62]

Griffiths-Jones S, Moxon S, Marshall M. et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 2005; 33:D121-4

[63]

Lowe TM, Eddy SR. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Res. 1997; 25:955-64

[64]

Lagesen K, Hallin P, Rødland EA. et al. RNAmmer: consistent and rapid annotation of ribosomal RNA genes. Nucleic Acids Res. 2007; 35:3100-8

[65]

Kent WJ. BLAT-the BLAST-like alignment tool. Genome Res. 2002; 12:656-64

[66]

Birney E, Clamp M, Durbin R. GeneWise and Genomewise. Genome Res. 2004; 14:988-95

[67]

Xu Z, Wang H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 2007; 35:W265-8

[68]

Price AL, Jones NC, Pevzner PA. De novo identification of repeat families in large genomes. Bioinformatics. 2005; 21:i351-8

[69]

Hoede C, Arnoux S, Moisset M. et al. PASTEC: an automatic transposable element classification tool. PLoS One. 2014; 9:e91929

[70]

Jurka J, Kapitonov VV, Pavlicek A. et al. Repbase update, a database of eukaryotic repetitive elements. Cytogenet Genome Res. 2005; 110:462-7

[71]

Tarailo-Graovac M, Chen NS. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinformatics. 2009; 4:4.10.11-14.10.14

[72]

Emms DM, Kelly S. OrthoFinder: phylogenetic orthology inference for comparative genomics. Genome Biol. 2019; 20:238

[73]

Mi HY, Muruganujan A, Ebert D. et al. PANTHER version 14: more genomes, a new PANTHER GO-slim and improvements in enrichment analysis tools. Nucleic Acids Res. 2019; 47:D419-26

[74]

Yu GC, Wang LG, Han YY. et al. ClusterProfiler: an R package for comparing biological themes among gene clusters. OMICS. 2012; 16:284-7

[75]

Katoh K, Asimenos G, Toh H. Multiple alignment of DNA sequences with MAFFT. Methods Mol Biol. 2009; 537:39-64

[76]

Suyama M, Torrents D, Bork P. PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments. Nucleic Acids Res. 2006; 34:W609-12

[77]

Talavera G, Castresana J. Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments. Syst Biol. 2007; 56:564-77

[78]

Kalyaanamoorthy S, Minh BQ, Wong TKF. et al. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods. 2017; 14:587-9

[79]

Nguyen LT, Schmidt HA, Von Haeseler A. et al. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol Biol Evol. 2015; 32:268-74

[80]

Yang ZH. PAML: a program package for phylogenetic analysis by maximum likelihood. Bioinformatics. 1997; 13:555-6

[81]

Puttick MN. MCMCtreeR: functions to prepare MCMCtree analyses and visualize posterior ages on trees. Bioinformatics. 2019; 35:5321-2

[82]

Han MV, Thomas GW, Lugo-Martinez J. et al. Estimating gene gain and loss rates in the presence of error in genome assembly and annotation using CAFE 3. Mol Biol Evol. 2013; 30:1987-97

[83]

Buchfink B, Xie C, Huson DH. Fast and sensitive protein alignment using DIAMOND. Nat Methods. 2015; 12:59-60

[84]

Wang YP, Tang H, DeBarry JD. et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res. 2012; 40:e49

[85]

Tang HB, Krishnakumar VJ, Li JP. jcvi: JCVI utility libraries. Zenodo. 2015

[86]

Xu YQ, Bi C, Wu G. et al. VGSC: a web-based vector graph toolkit of genome synteny and collinearity. Biomed Res Int. 2016; 2016:7823429

[87]

Zwaenepoel A, Van de Peer Y. Wgd-simple command line tools for the analysis of ancient whole-genome duplications. Bioinformatics. 2019; 35:2153-5

[88]

Rice P, Longden I, Bleasby A. EMBOSS: the European molecular biology open software suite. Trends Genet. 2000; 16:276-7

[89]

Ossowski S, Schneeberger K, Lucas-Lledó JI. et al. The rate and molecular spectrum of spontaneous mutations in Arabidopsis thaliana. Science. 2010; 327:92-4

[90]

Langfelder P, Horvath S. WGCNA: an R package for weighted gene co-expression network analysis. BMC Bioinformatics. 2008; 9:559

[91]

Shannon P, Markiel A, Ozier O. et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003; 13:2498-504

[92]

Prism 9 user guide, https://www.graphpad.com/guides/prism/latest/user-guide/index.htm.

[93]

Chen X, Zhu Q, Nie Y. et al. Determination of conifer age biomarker DAL1 interactome using Y2H-seq. For Res. 2021; 1:12

[94]

Chen CJ, Chen H, Zhang Y. et al. TBtools: an integrative toolkit developed for interactive analyses of big biological data. Mol Plant. 2020; 13:1194-202

[95]

Edgar RC. MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res. 2004; 32:1792-7

[96]

Capella-Gutiérrez S, Silla-Martínez JM, Gabaldón T. TrimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics. 2009; 25:1972-3

[97]

Stolzer M, Lai H, Xu M. et al. Inferring duplications, losses, transfers and incomplete lineage sorting with nonbinary species trees. Bioinformatics. 2012; 28:i409-15

[98]

Chen K, Durand D, Farach-Colton M. NOTUNG: a program for dating gene duplications and optimizing gene family trees. J Comput Biol. 2000; 7:429-47

PDF (2195KB)

128

Accesses

0

Citation

Detail

Sections
Recommended

/