INTRODUCTION
DNA sequencing technology has played a pivotal role in the advancement of molecular biology (
Gilbert, 1980). Over the past decade, we have witnessed tremendous transformation in this field. The neck-breaking speed of this change has been comparable with the rapid evolution of semiconductor industry under the Moore’s law (
Moore, 1965;
Shendure et al., 2004). The advancement of disruptive sequencing technologies afford us unprecedented high-throughput and low-cost sequencing platforms. The fast and low-cost sequencing approaches not only change the landscape of genomes sequencing projects but also usher in new opportunities for sequencing in various applications—the innovative ways of applying these technologies being created. Over the past few years, we have seen next-generation technologies being applied in a variety of contexts, including
de novo whole-genome sequencing, resequencing of genomes for variations, profiling mRNAs and other small and non-coding RNAs, assessing DNA binding proteins and chromatin structures, and detecting methylation patterns. Traditional and established techniques such as microarray in functional genomics applications have been increasingly challenged by the high-throughput next-generation sequencing techniques. With so many applications and instrument platforms available on the market, a common question would be: which platform is the best choice for a given biologic experiment? As a follow-up to our previous review article on generations of sequencing technologies (
Zhou et al., 2010), here we describe a few widely-used, commercially available next-generation instrument platforms and devote most of the following space to highlight their transforming potential and suitability to various applications. We believe that the synergetic relationship between sequencing technologies and its applications will ensure the continuing push toward faster, cheaper and more reliable approaches of producing sequences into the foreseeable future.
CURRENT NEXT-GENERATION SEQUENCING PLATFORMS
Even though the capillary sequencer that is based on Sanger’s chain-termination chemistry and developed in the mid-1970s (
Sanger and Coulson, 1975), such as the ABI 3730 (Applied Biosystems), is still viable and good for sequencing PCR products and other small-scale sequencing projects (
Mardis, 2008), most sequencing operations today are performed on next-generation instruments. The three dominate commercial platforms currently on the market are the Roche 454 Genome Sequencer, the Illumina Genome Analyzer, and the Life Technologies SOLiD System. To help readers understand the advantages and limitations of these instruments in relating to wide range of applications, a brief overview of their inner workings is in order.
All three platforms were developed at the end of 1990s and commercialized around 2005. They all adopted conceptually similar work flow as outlined in our previous article (
Zhou et al., 2010), from template library preparation, template amplification, to parallel sequencing by chain-extension, with variations in array formation, cluster generation and enzyme-based sequencing biochemistry. New breed of sequencing-by-synthesis instruments with single molecule detection is also on the horizon. Recently, Helicos Biosciences has introduced its version of single-molecule sequencing (tSMS), the Helicos Genetic Analysis System (
http://www.helicosbio.com/). Pacific Bioscience also introduced its Zero-Mode Waveguide based SMRT technology (
http://www.pacificbiosciences.com). Comparison of the next-generation sequencing platforms is summarized in Table 1.
The Roche 454 Genome Sequencer FLX system
The GS FLX system based on sequencing-by-synthesis with pyrophosphate chemistry, was developed by 454 Life Sciences and was the first next-generation sequencing platform available on the market (
Margulies et al., 2005). In this system, the DNA sample is first sheared into fragments. Two short adaptors, an A-adaptor and a B-adaptor are then ligated to the fragments. The adaptors provide priming sites for amplification and sequencing, as well as a special key sequence. The B-adaptor also contains a 5ʹ-biotin tag that enables the immobilization of library fragments onto streptavidin-coated magnetic beads. The double-stranded products bound to the beads are then denatured to release the complementary non-biotinylated strands containing both an A- adaptor sequence and a B-adaptor sequence. These denatured strands form the single-stranded template DNA library (Fig. 1A). For DNA amplification, the Genome Sequencer FLX system employs emulsion-based clonal amplification, called emPCR (
Dressman, 2003). The single-stranded DNA library is immobilized by hybridization onto primer-coated capture beads. The process is optimized to produce beads where a single library fragment is bound to each bead. The bead-bound library is emulsified along with the amplification reagents in a water-in-oil mixture. Each bead with a single library fragment is captured within its own emulsion microreactor, where the independent clonal amplification takes place. After amplification, the microreactors are broken, releasing the DNA-positive beads for further enrichment (Fig. 1B). For sequencing, the DNA beads are layered onto a PicoTiterPlate device, depositing the beads into the wells, followed by enzyme beads and packing beads. The enzyme beads contain sulfurylase and luciferase, which are key components of the sequencing reaction, while the packing beads ensure that the DNA beads remain positioned in the wells during that sequencing reaction (Fig. 1C). The fluidics sub-system delivers sequencing reagents that contain buffers and nucleotides by flowing them across the wells of the plate. Nucleotides are flowed sequentially in a specific order over the PicoTiterPlate device. When a nucleotide is complementary to the next base of the template strand, it is incorporated into the growing DNA strand by the polymerase. The incorporation of a nucleotide releases a pyrophosphate moiety. The sulfurylase enzyme converts the pyrophosphate molecule into ATP using adenosine phosphosulfate. The ATP is hydrolyzed by the luciferase enzyme using luciferin to produce oxyluciferin and give off light. The light emission is detected by a CCD camera, which is coupled to the PicoTiterPlate device. The intensity of light from a particular well indicates the incorporation of nucleotides (Fig. 1D). Across multiple cycles, the pattern of detected incorporation events reveals the sequence of templates represented by individual beads. The sequencing is ‘asynchronous’ in that some features may get ahead or behind other features depending on their sequence relative to the order of base addition. Raw reads processed by the 454 platform are screened by various quality filters to remove poor-quality sequences, mixed sequences (more than one initial DNA fragment per bead), and sequences without the initiating key sequence. For downstream analysis, three different bioinformatic tools are available: GS
De Novo Assembler, GS Reference Mapper, and GS Amplicon Variant Analyzer (
http://454.com/products-solutions/analysis-tools/index.asp). Using these graphical analysis tools, researchers can quickly obtain biologically informative results from sequence data.
A major limitation of the 454 technology relates to resolution of homopolymer-containing DNA segments, such as AAA and GGG (
Rothberg and Leamon, 2008). Because there is no terminating moiety preventing multiple consecutive incorporations at a given cycle, pyrosequencing relies on the magnitude of light emitted to determine the number of repetitive bases. This is prone to a greater error rate than the discrimination of incorporation versus nonincorporation. As a consequence, the dominant error type for the 454 platform is insertion-deletion, rather than substitution. Another disadvantage of 454 sequencing platform is that the per-base cost of sequencing is much higher than that of other next-generation platforms, e.g., SOLiD and Solexa (
Rothberg and Leamon, 2008). It is therefore unsuitable for sequencing targeted fragments from small numbers of DNA samples, such as those for phylogenetic analysis. Comparing to other next-generation platforms, the key advantage of the 454 platform is its read length (
Metzker, 2010). The 454 System can generate more than 1,000,000 individual reads with improved Q20 read length of 400 bases per 10 h instrument run. It may be a best choice for certain applications where long read-lengths are critical, such as
de novo assembly and metagenomics. The company is also projected to place a low-throughput version (1/10 of the GS throughput) of the instrument GS Junior, into the market later this year, which is ideal for sequencing bacterial genomes and yields about 20 × coverage of a typical bacterial genome, ~5 Mb in size (
http://www.gsjunior.com/).
The Illumina (Solexa) Genome Analyzer
The Solexa sequencing platform was commercialized in 2006. The working principle (Fig. 2) is sequencing-by-synthesis chemistry. Input DNA is fragmented by hydrodynamic shearing to generate < 800 bp fragments. The fragments are blunt ended and phosphorylated, and a single ‘A’ nucleotide is added to the 3ʹ-ends of the fragments. Then DNA fragments are ligated at both ends to adapters that have a single-base ‘T’ overhang. After denaturation, DNA fragments are immobilized at one end on a solid support-flow cell. The surface of the flow cell is coated densely with the adapters and the complementary adapters. Each single-stranded fragment that is immobilized at one end on the surface creates a ‘bridge’ structure by hybridizing with its free end to the complementary adapter on the surface of the flow cell. The adapters on the surface also act as primers for the following PCR amplification. Adding mixtures containing the PCR amplification reagents to the flow cell surface, the DNA fragments are amplified by “bridge PCR” (
Adessi, 2000;
Fedurco et al., 2006). After several PCR cycles, about 1000 copies of single-stranded DNA fragments are created on the surface, forming a surface-bound colony (the cluster). The reaction mixture for the sequencing chemistry and DNA synthesis is supplied onto the surface, which contains four reversible terminator nucleotides, each labeled with a different fluorescent dye. After incorporation into the DNA strand, the terminator nucleotide as well as its position on the support surface are detected and identified via its fluorescent dye by the CCD camera. The terminator group at the 3ʹ-end of the base and the fluorescent dye are then removed from the base and the synthesis cycle is repeated. This series of steps continues for a specific number of cycles, as determined by user-defined instrument settings. A base-calling algorithm assigns sequences and associated quality values to each read and a quality checking pipeline evaluates the Illumina data from each run, removing poor-quality sequences.
In 2008, Illumina introduced an upgrade, the Genome Analyzer II, to its predecessor, which offered a powerful combination of the cBot and Paired-End Module (
http://www.illumina.com/systems/genome_analyzer.ilmn). cBot is a revolutionary automated system that creates clonal clusters from single molecule DNA templates, preparing them for sequencing by synthesis on the Genome Analyzer. The Paired-End Module is a fluidics station that attaches to the Genome Analyzer. After completion of the first read, the templates can be regenerated
in situ to prepare for the second round of sequencing from the opposite end of the fragments. First, the newly sequenced strands are stripped off and the complementary strands are bridge amplified to form clusters. Once the original templates are cleaved and removed, the reverse strands undergo sequencing by synthesis. The Paired-End Module enables paired-end sequencing up to 2 × 100 bp for fragments ranging from 200 bp to 5 kb. For Genome Analyzer II, the run time is highly decreased and the output per paired-end run can reach 45–50 Gb (gigabasepairs). Compared to Sanger sequencing, the Illumina system is able to produce more data at a reduced time and cost; however, error rates are higher (often resulting in false-positive when identifying sequence variations) and reads are shorter (
Metzker, 2010). Usually error rate can be overcomed by coverage but contiguity is rather limited by the read length as the Lander-Waterman Curve describes (
Lander and Waterman, 1988). Illumina also offers a newer version of its next-generation sequencer, HiSeq2000, which has a dual flowcell system installed to increase the efficiency of data generation (
http://www.illumina.com/systems/hiseq_2000.ilmn).
The Life Technologies SOLiD system
The Life Technologies SOLiD system is based on a sequencing-by-ligation technology. This platform has its origins in the system described by
Shendure et al. (2005) and in work by
McKernan and colleagues (2006) at Agencourt Personal Genomics (acquired by Applied Biosystems in 2006). The generation of a DNA fragment library and the sequencing process by subsequent ligation steps are shown in Fig. 3. In this technology, two types of libraries—fragment or mate-paired library can be constructed depending on the researchers’ purposes. Then DNA fragments are ligated to adapters and bound to beads. DNA fragments on the beads are amplified by the emulsion PCR, after which the templates are denatured and bead enrichment is performed to select beads with extended templates. The template on the selected beads undergoes a 3ʹ modification to allow covalent bonding to the glass slide. The sequencing methodology is based on sequential ligation with dye-labeled oligonucleotides (
Housby and Southern, 1998). In the first step, a primer is hybridized to the adapter sequence within the library template. Next, a set of four fluorescently labeled oligonucleotide octamers compete for ligation to the sequencing primer. In these octamers, the first and second di-base are characterized by one of four fluorescent labels at the end of the octamer. After the detection of the fluorescence from the label, bases 1 and 2 in the sequence are thus determined. The ligated octamer oligonucleotides are cleaved off after the fifth base, removing the fluorescent label, then hybridization and ligation cycles are repeated, this time determining bases 6 and 7 in the sequence; in the subsequent cycle bases 11 and 12 are determined, etc. Progressive rounds of octamer ligation enable sequencing of every five bases. Following a series of ligation cycles, the extension product is removed and the template is reset with another octamer complementary to the n-1 position for a second round of ligation cycles. After five rounds of primer reset completed for each sequence tag, each base is interrogated in two independent ligation reactions by two different primers. This method is called ‘Two Base Encoding’ (
Mckernan et al., 2006). Two Base Encoding is a unique and powerful approach designed to clearly discriminate measurement errors. The combination of ligase enzymology, primer reset and two-base encoding all contribute to the low error rate and reduced systemic noise.
In December 2009, Applied Biosystems has updated the platform to SOLiD™ 3 Plus (
http://www3.appliedbiosystems.com/AB_Home/applicationstechnologies/SOLiD-System-Sequencing-B/index.htm). The SOLiD™ 3 Plus System can generate more than 60 Gb of mappable sequence or greater than 1 billion reads per run. The SOLiD™ 4 system is claimed to produce comparable amount of raw data as its competitors claimed. The cost of the instrument is substantially lower than that of other next-generation sequencing platforms. The current read length, however, significantly limits its applications (
Metzker, 2010).
Single molecule sequencing
The HeliScope Genetic Analysis System by Helicos Biosciences, based on the work from Quake’s group (
Braslavsky et al. 2003), is the first single molecule sequencing system and recently available on market. It utilizes sequencing by synthesis on single molecule. Constructed single-stranded DNA library is disorderly arrayed on a flat substrate without any amplification. DNA polymerase and one of four fluorescently labeled nucleotides are flowed into the system at each sequencing cycle. Strands in the array that have undergone template-directed base extension are lightened up by fluorescent label, which are recorded with a CCD camera. After washing, fluorescent labels on the extended strands are chemically removed, and another cycle of base extension repeats. Like in pyrosequencing with 454, each iterative cycle is asynchronous—some strands in the array may pull ahead, fall behind, or completely fail to extend all together, leading to similar problem with homopolymer run. However, unlike Roche 454 platform, single molecule affords us mitigation by playing trick with enzyme kinetics to slow down the rate of chain extension so to reduce the chance of two consecutive base incorporations before being washed away (
Harris et al. 2008).
One of the key challenges this technology faces is the raw sequencing accuracy due to the difficulty with detecting single molecule event. Therefore, the dominant error type with this instrument is deletion. A two-pass strategy can somewhat mitigate the shortcoming. Single molecule sequencing means that we can now reset the tethered template DNA to its original state by lifting off the newly extended strand after one sequencing run. Another sequencing pass can then be performed in opposite direction from distal adaptor, yielding a second sequence from the same template. The duplicated sequences can be used to average out detection errors, and thus, give rise to much higher accuracy than otherwise.
Pacific Biosciences is another company currently developing the single molecule sequencing technology, the SMRT (Single Molecule Real Time) technology (
Eid et al., 2009). The SMART technology relies on a nano-structure, the Zero Mode Waveguide (ZMW), for real-time observation of DNA polymerization (
Levene et al., 2003). ZMW chip consists of thousands upon thousands of sub-wavelength holes, tens of nanometers in diameter, fabricated by perforating a thin metal film supported by transparent substrate. When illuminated from the side of glass, light is not able to penetrate through the hole but leaves an exponentially-decayed evanescent wave at very bottom of each hole, and thus, creating a very small volume of fluorescence detection. Furthermore, the DNA polymerase is planted to the bottom of each waveguide. During a sequencing assay, each time a fluorescently labeled base is grabbed on by the polymerase, it brings fluorophore to the detection volume, creating a burst of fluorescent light. If the nucleotide is complementary to the template strand, it will go through a time-consuming synthesis process, and therefore, stay in the detection volume longer until the fluorescent moiety is released as part of pyrophosphate. The color-coded fluorescence burst and its duration reveal the identity of the complementary base on template DNA. By continuously following the bursts of fluorescence at each waveguide in real-time, sequences of template DNA can be rapidly determined.
This technology has the great potential to achieve high speed with long read length. However, error stemmed from real-time single-molecule detection might put a damper on its raw accuracy as with other SMS-based sequencing platforms. Current CCD technology has also limited the maximum ZMW chip area that can be simultaneously observed. Low yield ratios (~30%) of polymerase-occupied waveguide further limit the number of useable wells on the chip (
Korlach et al., 2008). Even with this limitation, the first version of the instrument when introduced has promised a read length of no less than 1500 bp, at a speed of 15 min per run, and with reagent cost no more than $60 per run. It is anticipated that future version of this platform, after technical issues are resolved, could churn out 100 Gb of data per day with read length up to 100,000 bp.
APPLICATION OF NEXT-GENERATION SEQUENCING TECHNOLOGIES
The production of low cost reads by next-generation sequencing technologies makes them useful in a variety of areas (Table 2). Important applications include: (1)
de novo genome sequencing, whole-genome resequencing or more targeted sequencing for discovery of mutations or polymorphisms; (2) transcriptome analysis and cataloguing, where shotgun libraries derived from mRNA or small RNAs are deeply sequenced; (3) large-scale analysis of DNA methylation, by deep sequencing of bisulfite-treated DNA; (
http://454.com/products-solutions/analysis-tools/index.asp) genome-wide mapping of DNA-protein interactions, by deep sequencing of DNA fragments pulled down by chromatin immunoprecipitation (ChIP-Seq); (
http://454.com/products-solutions/analysis-tools/index.asp) species classification and/or gene discovery by metagenomics and pangenomics. As mentioned previously, there are different advantages and limitations among the next-generation platforms in respect to specific applications.
De novo sequencing and assembly
De novo sequencing is the initial generation of the primary genomic sequence of a particular organism. A detailed genetic analysis of any organism is possible only after
de novo sequencing has been performed (
Goldberg et al., 2006;
Durfee et al., 2008;
Reinhardt et al., 2009), i.e., reference sequence produced. Until March 1, 2010, there have been 740 eukaryotic genome sequencing projects submitted to NCBI (Table 3), while only 23 genomes are completed, and most of them are in draft assemblies or work-in-progress (
http://www.illumina.com/systems/genome_analyzer.ilmn). Four of these organisms are sequenced using next-generation sequencing technologies independently or in combination with the traditional Sanger method (
Velasco et al., 2007;
Diguistini et al., 2009;
Huang et al., 2009;
Li et al., 2010) (Table 4). A hybrid Sanger/pyrosequencing approach resolved a complex heterozygous grape genome, where consensus sequence of the genome and a set of mapped marker loci were generated. This is the first project that utilizes both the long Sanger and short SBS reads to assemble the genome sequence of a large eukaryotic genome (
Velasco et al., 2007). A draft sequence of the giant panda genome was successfully generated and assembled based on next-generation sequencing technology alone, taking the advantage of excellent colinerarity of the mammalian genomes. The assembled contigs (2.25 Gb) cover approximately 94% of the whole genome and the remaining gaps (0.05 Gb) seem to contain carnivore-specific repeats and tandem repeats (
Li et al., 2010). When taking on large genomes, i.e., over 1 Gb in total length, one should be more cautious in designing a sequencing experiment since the effort could be severely hampered by polyploidy and large repetitive fraction. Nevertheless, successful sequencing projects demonstrate the feasibility of using next-generation sequencing technologies for accurate, cost-effective, and rapid
de novo assembly of large eukaryotic genomes (
Imelfort and Edwards, 2009;
Turner et al., 2009a).
With its long read lengths and high accuracy, capillary electrophoresis-based sequencing has been the gold standard for de novo genome sequencing projects in the past decades. However, the throughput of these systems makes de novo assembly of most organisms a lengthy and costly endeavor. Next-generation sequencing technologies hold great promise in reducing the time and cost. Compared to just a few years ago, it is now much easier and cheaper to sequence entire genomes, and a wide variety of species are being studied using these advanced tools every day.
Whole-genome or targeted resequencing
By far, the most common use of next-generation sequencing platform has been resequencing (
Davies, 2007). To identify single nucleotide polymorphisms, indels, copy number and structural variations, multiple individuals or strains, or a population-based sampling of a species have to be resequenced (
Bentley, 2006;
Ossowski et al., 2008;
Denver et al., 2009;
Xia et al., 2009;
Pleasance et al., 2010). In humans, such an endeavor has already commenced with the publication of several complete genomes (Table 5), with the list growing by the day. The first is from J. Craig Venter and achieved using traditional Sanger sequencing methods (
Levy et al., 2007) as part of the Human Genome Project and the second is from James D. Watson, which was sequenced using the Roche 454 technology to 7.5 × genome coverage. The reads were aligned to the NCBI reference sequence using a combination of the BLAT and Smith-Waterman algorithms. The sequence differs from the reference at 3.32 Mb, of which 2.7 Mb are known differences (
Wheeler et al., 2008). The next four human genome sequences are from a Chinese (
Wang et al., 2008), an African (
Pushkarev et al., 2009), and two Korean individuals (
Ahn et al., 2009;
Kim et al., 2009); all were done using the Illumina Genome Analyzer and sequenced to around 20 × haploid genome coverage with the exception of the African male’s genome which was also resequenced on ABI SOLiD system (
McKernan et al. 2009). For all four genomes, reads covered more than 99% of the NCBI human reference genome, revealing approximately 3 million SNPs. More recently, James Lupski’s genome was sequenced to 30 × base coverage using ABI’s SOLiD System (
Lupski et al., 2010). Resequencing of human genome was not limited to the 2nd-generation platforms. Steven Quake’s genome, for example, was sequenced to 90% genome coverage on Helicos’ single-molecule sequencing platform (
Pushkarev et al., 2009).
Target-region resequencing refers to sequencing a targeted region of a species’ genome from multiple individuals; it enables scientists to investigate variations of interested genomic regions or genes with high coverage and lower cost (
Harismendy et al., 2009). Two methods of target-region resequencing are widely used: PCR-based candidate gene (
Dracatos et al., 2009;
Goossens et al., 2009;
Harismendy and Frazer, 2009;
Tewhey et al., 2009) and whole exome approaches (
Hodges et al., 2007;
Porreca et al., 2007;
Choi et al., 2009;
Turner et al., 2009b). Ji and colleagues developed a procedure for massive parallel resequencing of multiple human genes. It combines a highly multiplexed and target-specific amplification process with a parallel sequencing technology (
Dahl et al., 2007). They demonstrated parallel resequencing of 10 cancer genes covering 177 exons with average sequence coverage per sample of 93%. Through exome sequencing, Bamshad et al. discovered the gene for a rare mendelian disorder of unknown cause, the Miller syndrome (
Ng et al., 2010), which demonstrates that exome sequencing of a small number of unrelated affected individuals is a powerful, efficient strategy for identifying the genes underlying rare mendelian disorders and will likely transform the genetic analysis of monogenic traits.
Whole transcriptome shotgun sequencing: RNA-Seq
The transcriptome is the complete set of transcripts in a cell, and their quantity, for a specific developmental stage or physiologic condition (
Jacquier, 2009). Understanding the transcriptome is essential for interpreting the functional elements of the genome and revealing the molecular constituents of cells and tissues, and also for understanding development and disease. The specific aims of transcriptomics are: (1) to catalog all transcripts in a context of cell types for a species, including mRNAs, non-coding RNAs and small RNAs; (2) to determine the transcriptional structure of genes, in terms of their start sites, 5ʹ- and 3ʹ-ends, splicing patterns and other post-transcriptional modifications; and (3) to quantify the expression levels of each transcript during development or under different physiologic and pathological conditions. With the availability of faster and cheaper next-generation sequencing platforms, more transcriptomic analyses are performed using a recently-developed deep sequencing approach, RNA-Seq (
Wang et al., 2009). Studies using this method have already altered our view of the extent and complexity of eukaryotic transcriptomes (
Cloonan et al., 2008;
Mortazavi et al., 2008;
Sugarbaker et al., 2008;
Sultan et al., 2008;
Tang et al.,2010).
The current gold standard for protein-coding gene annotation is EST or full-length cDNA sequencing followed by alignment to a reference genome, but it has been estimated that most EST studies using Sanger sequencing detect only about 60% of transcripts in the cell, which fails to cover the poorly expressed or long transcripts (
Brent, 2008). This information gap can be addressed using the next-generation sequencing technologies, which have been used to generate transcriptomes for many species and tissues (
Mortazavi et al., 2008;
Nagalakshmi et al., 2008;
Sultan et al., 2008). For instance, a study used the 454 technology to generate 391,157 EST reads from the brain transcriptome of the wasp
P. metricus (
Toth et al., 2007). The reads were then aligned to the genome sequence and EST resources from the honeybee,
Apis mellifera, to annotate
P. metricus transcripts. Interestingly, the study found wasp EST matches to 39% of the honeybee mRNAs and observed a strong correlation between the expression levels of the corresponding transcripts from the two species.
The short reads produced by high-throughput next-generation technologies, particularly Illumina and SOLiD, are arguably suitable for gene expression profiling based on tens of millions of short reads rather than tens of thousands of based on the Sanger method. RNA-Seq has been used to accurately monitor gene expression during yeast vegetative growth (
Nagalakshmi et al., 2008), yeast meiosis (
Wilhelm et al., 2008) and mouse embryonic stem-cell differentiation (
Cloonan et al., 2008), to track gene expression changes during development, and to provide a ‘digital measurement’ of gene expression difference among different tissues.
Before the advent of transcriptome shotgun sequencing, the starts and ends of most transcripts had not been precisely resolved and the extent of spliced heterogeneity remained poorly understood. RNA-Seq, with its high resolution and sensitivity, has revealed many novel transcribed regions and splicing isoforms of known genes. It also helps to map 5ʹ- and 3ʹ-boundaries of many genes. Using RNA-seq method, the 5ʹ- and 3ʹ-boundaries of 80% and 85% of all annotated genes, respectively, were mapped in
S. cerevisiae (
Nagalakshmi et al., 2008). Similarly, in
S. pombe (
Wilhelm et al., 2008) many boundaries were defined by RNA-Seq data in combination with tiling array data. In humans, 31,618 known splicing events were confirmed (11% of all known splicing events) and 379 novel splicing events were discovered (
Morin et al., 2008a). In mice, extensive alternative splicing was observed for 3462 genes (
Mortazavi et al., 2008). In addition, results from RNA-Seq suggest the existence of a large number of novel transcribed regions in every genome surveyed, including those of
A. thaliana (
Lister et al., 2008), mouse (
Cloonan et al., 2008;
Mortazavi et al., 2008), human (
Morin et al., 2008a),
S. cerevisiae (
Nagalakshmi et al., 2008) and
S. pombe (
Wilhelm et al., 2008). These novel transcribed regions, combined with many undiscovered novel splicing variants, suggest that there is considerably more transcriptomic complexity than previously appreciated.
Small RNA analysis
A related application of next-generation sequencing technologies to the analysis of transcriptomes is small RNA discovery and profiling. High-throughput sequencing offers a greater potential for the identification of novel small RNAs as well as profiling of known and novel small RNA genes. Small RNA profiling with 454 pyrosequencing technology has been widely reported, which include studies in the moss
Physcomitrella patens(
Axtell et al., 2006),
A. thaliana (
Henderson et al., 2006;
Lu et al., 2006;
Rajagopalan et al., 2006),
Triticum aestivum (
Yao et al., 2007), the basal eudicot species
Eschscholzia californica (
Barakat et al., 2007), the lycopod
Selaginella moellendorffii (
Axtell et al., 2006), the unicellular alga
Chlamydomonas reinhardtii (
Zhao et al., 2007), Marek disease virus (
Burnside et al., 2006), and some primates (
Berezikov et al., 2006). More importantly, these studies contributed to the discovery of a novel class of small RNAs, termed Piwi-interacting RNAs. They are expressed in mammalian testes and are presumably required for germ cell development in mammals and other species (
Girard et al., 2006;
Lau et al., 2006;
Houwing et al., 2007).
The higher throughput of Illumina and SOLiD technologies enables the generation of deeper small RNA libraries. Using Illumina sequencing,
Morin et al. (2008b) identified 334 known plus 104 novel miRNA genes expressed in human embryonic stem cells, while
Glazov et al. (2008) detected 449 novel and all known chicken miRNAs in the chicken embryo. In addition, small RNA profilings in locust,
Xenopus tropicalis (
Armisen et al., 2009),
C. elegans embryos (
Stoeckius et al., 2009) and
Gossypium hirsutum L (
Pang et al., 2009) have also been reported.
Epigenomic Analysis
Epigenetics is the study of heritable gene regulation that does not involve the DNA sequence itself but its modifications and higher-order structures. The next-generation sequencing technologies offer the potential to substantially accelerate epigenomic research. To date, these technologies have been applied in several epigenomic areas, including the characterization of DNA methylation patterns, posttranslational modifications of histones, the interaction between transcription factors and their direct targets, and nucleosome positioning on a genome-wide scale. These areas are summarized into the following two major sections.
Methylome
DNA cytosine methylation is a central epigenetic modification that plays essential roles in cellular processes including genome regulation, development and disease. Single-base resolution analysis of DNA methylation sites can be achieved by sodium bisulfite (BS) treatment of genomic DNA, which converts cytosines, but not methylcytosines, to uracil. Subsequent sequencing of PCR-amplified bisulfite-converted DNA allows determination of the methylation state of the cytosines in the sequenced region of the genome, as methylcytosine will be sequenced as cytosine, and unmethylated cytosine as thymine.
Taylor et al. (2007) improved the bisulfite DNA sequencing procedure by combining with the 454 sequencing technology. The approach was applied to analyze methylation patterns in 25 gene-related CpG-rich regions from > 40 cases of primary cells. The study generated > 1600 individual sequence far beyond the few clones ( < 20) typically analyzed by traditional bisulfite sequencing. Using the Illumina Genome Analyzer,
Cokus et al. (2008) and
Lister et al. (2008) generated 2–3 gigabases of uniquely aligned bisulfite sequence to comprehensively identify sites of DNA methylation throughout the
Arabidopsis genome at a single-base resolution, including previously unidentified sites of cytosine methylation, and local sequence motifs associated with DNA methylation. The approach of bisulfite DNA sequencing is widely used for DNA methylation profiling in various organisms now (
Costello et al., 2009;
Smith et al., 2009;
Bormann Chung et al., 2010).
DNA-protein interactions: ChIP-Seq
ChIP-seq is a recently developed technique for genome-wide profiling of DNA binding proteins, histone modifications and nucleosomes. The association between DNA and proteins is a fundamental biologic interaction that plays a key part in regulating gene expression and controlling the availability of DNA for transcription, replication and other biologic processes. These interactions can be studied using a technique called chromatin immunoprecipitation, and traditionally followed by microarray analysis (
Aparicio et al., 2004). With the advent of high-throughput next-generation sequencing technologies, a more powerful approach based on chromatin immunoprecipitation followed by sequencing (ChIP-seq) has emerged. The precedent-setting paper for ChIP-seq was published by Johnson and colleagues in 2006, who used
Caenorhabditis elegans and the Roche 454 platform to elucidate nucleosome positioning on genomic DNA (
Johnson et al., 2006). This study established that sequencing the nuclease-derived digestion products of genomic DNA was sufficient to generate a genome-wide, highly precise positional profile of chromatin. Subsequent studies utilized a ChIP-based approach and the Illumina platform to provide insights into transcription factor binding sites in the human genome such as neuron-restrictive silencer factor (NRSF) (
Johnson et al., 2007) and signal transducer and activator of transcription 1 (STAT1) (
Robertson et al., 2007). The first applications of ChIP-seq to profile histone modifications were done in CD4+ T cells (
Impey et al., 2004) and mouse embryonic stem (ES) cells (
Mikkelsen et al., 2007). In a landmark study, Mikkelsen and coworkers (2007) explored the connection between chromatin packaging of DNA and differential gene expression using mouse embryonic stem cells and lineage-committed mouse cells (neural progenitor cells and embryonic fibroblasts). Their work provided a next-generation sequencing-based framework for using genome-wide chromatin profiling to characterize cell populations.
In comparison to array-based predecessor, i.e., ChIP-chip technology, ChIP-seq offers higher resolution, lower noise and better coverage. With the ever-decreasing cost of sequencing, ChIP-seq has become an indispensable tool for studying gene regulation and epigenetic mechanisms.
Metagenomic sequencing
Metagenomics involves the genomic analysis of microorganisms by direct extraction of DNA from uncultured ensemble of microbial communities. It is not until recent years that scientists are able to unravel a wider range of microorganisms, thanks largely to advances in DNA-sequencing technology. Robert Edwards and colleagues published the first sequences of environmental samples generated with next generation sequencing technique with Roche’s 454 pyrosequencing instrument (
Edwards et al., 2006). Since then, a wide range of metagenomes has been studied with this technique, including some real large metagenomic project such as the Human Microbiome Project (HMP) (
Turnbaugh et al., 2007). The aim of HMP is to lay bare the microbial communities associated with various parts of the human body, including the gut. In a recent publication, an international team of researchers have cataloged human gut microbial genes by metagenomic sequencing (
Qin et al., 2010). They generated over 570 Gb of sequence data from 124 individuals, assembled and characterized 3.3 million non-redundant microbial genes. This helped scientists, for the first time, to define the minimal human gut metagenome and the minimal gut bacterial genomes. Even though it has been widely accepted that the 454 system is the most promising next-generation sequencing technology for metagenomic analysis, due to its long read length, this recent work on human gut microbial was done with Illumina Genomic Analyzer.
CONCLUSIONS
The field of sequencing technology and application development is a fast-moving area of biomedical research. As we can see from the introduction above, the next-generation sequencing technologies have extended to an impressive array of applications beyond just genomic sequencing and its large-scale operations. As we described, new innovations are being developed each day. Novel generations of sequencing technologies, such as single-molecule sequencing and nanostructure-based sequencing, holds greater promise to achieve ever-faster, cheaper, more accurate and reliable ways to produce sequence data. Shortcomings of today’s next-generation sequencing platforms, e.g., short-read and less base accuracy, will be overcomed with the development of new technologies. This, indeed, makes this an exciting area for genomic studies.
Higher Education Press and Springer-Verlag Berlin Heidelberg 2010