Skip to main content
ExplainerGenomic SequencingExplainer· 4 min read· in Science

Decoding the Genome: How Synthesis, SMRT, and Nanopore Architectures Read DNA

Three distinct chemical architectures dominate modern DNA sequencing, trading off between extreme accuracy, continuous read length, and direct epigenetic detection. Understanding how each platform physically reads nucleotide bases reveals why no single technology can map the entire human genome alone.

By Viktoria Sokolova

Short-Read Advocates 40%Long-Read Specialists 35%Direct-Detection Pioneers 25%
Short-Read Advocates
Prioritize extreme accuracy, massive throughput, and cost-efficiency for population-scale genomics and targeted oncology panels.
Long-Read Specialists
Argue that circular consensus sequencing is necessary to resolve complex structural variants and repetitive regions without sacrificing accuracy.
Direct-Detection Pioneers
Focus on the biophysical advantages of reading native DNA, emphasizing ultra-long continuous reads and simultaneous epigenetic profiling.

Perspectives this story doesn't cover

  • Clinical Geneticists
  • Bioinformatics Software Developers

What we don’t know

  • The exact physical limit of nanopore base-calling accuracy, and whether neural networks can eventually push it to parity with synthesis methods.
  • How quickly the cost of long-read sequencing will drop to match the current $100-per-genome standard set by short-read platforms.
  • The full extent of structural variations hidden in the human population that short-read sequencing has historically failed to detect.

The moment a sequencing platform commits to a base call—identifying an adenine, cytosine, guanine, or thymine—dictates the entire trajectory of genomic analysis. Whether that call relies on a fluorescent flash, a polymerase enzyme's hesitation, or a microscopic shift in electrical current determines not just the accuracy of the read, but what kind of biological questions the data can answer.[8]

Illumina's architecture, known as sequencing by synthesis (SBS), relies on optical imaging. As a DNA strand is copied, the system introduces nucleotides equipped with reversible terminators and fluorescent tags. When a base is incorporated, the synthesis temporarily halts, a laser excites the tag, and a camera records the color. The terminator is then cleaved, and the cycle repeats. This method produces short reads, typically 150 to 300 base pairs in length, but achieves an accuracy rate exceeding 99.9 percent.[1][3]

The limitation of the SBS approach lies in its read length. The human genome contains highly repetitive regions—stretches where the same sequence repeats thousands of times. Reconstructing these regions from 300-base-pair fragments is mathematically analogous to assembling a jigsaw puzzle depicting a cloudless blue sky. The pieces fit together perfectly, but their true position remains ambiguous, leaving structural variants and large insertions hidden from clinical view.[3][5]

Read lengths vary by orders of magnitude depending on the underlying sequencing chemistry.

To solve the assembly problem, Pacific Biosciences (PacBio) engineered Single Molecule, Real-Time (SMRT) sequencing. Instead of halting synthesis to take a photograph, SMRT sequencing anchors a single polymerase enzyme at the bottom of a microscopic well called a zero-mode waveguide (ZMW). As the enzyme naturally incorporates fluorescently labeled bases into a continuous DNA strand, the ZMW illuminates only the exact moment of incorporation, recording a real-time movie of the synthesis.[2][7]

In 2019, PacBio introduced HiFi sequencing, which fundamentally altered the accuracy landscape for long reads. By attaching hairpin adapters to both ends of a double-stranded DNA fragment, the molecule is turned into a closed loop. The polymerase travels around this loop repeatedly, reading the same 10,000 to 25,000 base-pair sequence multiple times. "HiFi sequencing provides both the high accuracy of short reads and the long read lengths required to assemble complex genomes," PacBio's 2024 technical documentation states, noting that the consensus of these multiple passes achieves 99.9 percent accuracy.[2][7]

In 2019, PacBio introduced HiFi sequencing, which fundamentally altered the accuracy landscape for long reads.

Oxford Nanopore Technologies discards synthesis entirely, relying instead on biophysics. A protein nanopore is embedded in an electrically resistant synthetic membrane. A motor protein ratchets a single strand of DNA through the pore at a controlled speed. As the DNA passes through, it physically obstructs the flow of ions across the membrane, creating a disruption in the electrical current.[4]

Because each combination of nucleotides has a distinct physical shape, it creates a unique electrical signature. Algorithms translate these current fluctuations back into base calls. This architecture places no chemical limit on read length; if a DNA molecule is extracted intact, the pore will read it. Nanopore platforms have successfully generated continuous reads exceeding 2 million base pairs, allowing researchers to span entire centromeres in a single pass.[4][8]

The trade-off for nanopore sequencing has historically been accuracy. Because the electrical signal is influenced by multiple bases residing in the pore simultaneously, base-calling algorithms must untangle complex overlapping signals. While early iterations struggled with error rates above 10 percent, modern neural-network-driven base callers have pushed raw accuracy into the 95 to 99 percent range, though it still trails the strict 99.9 percent standard of synthesis-based methods.[4][5]

The historical trade-off between continuous read length and raw base-calling accuracy.

However, nanopore's direct physical measurement offers a unique clinical advantage: the detection of epigenetics. Chemical modifications to DNA, such as methylation, alter the physical shape of the nucleotide and therefore change its electrical signature. A recent MDPI evaluation comparing Illumina and Nanopore platforms for studying Alzheimer's disease and frontotemporal dementia highlighted this capability.[6]

The MDPI researchers found that while Illumina requires harsh bisulfite conversion to detect methylation—a process that degrades DNA and complicates library preparation—nanopore sequencing reads the methylation marks directly from the native DNA strand. This allows a single sequencing run to capture both the genetic sequence and the epigenetic modifications simultaneously, preserving the biological context of the sample.[6]

Synthesis relies on optical imaging of fluorescent tags, while nanopore technology measures physical disruptions in electrical current.

The evidence indicates that no single platform currently satisfies every genomic requirement. Illumina remains the standard for high-throughput, low-cost variant calling in oncology. PacBio HiFi is the architecture of choice for resolving complex structural variations and generating reference-quality genomes. Oxford Nanopore dominates rapid pathogen surveillance, ultra-long scaffolding, and direct epigenetic profiling.[3][5][8]

The next threshold in genomic chemistry is not a single universal sequencer, but the seamless computational integration of all three signals. When a single software pipeline can simultaneously process the fluorescent flashes of synthesis, the timing of a trapped polymerase, and the electrical resistance of a protein pore, the genome will finally be read exactly as it functions.[8]

Key points

  1. Illumina's sequencing by synthesis provides extreme accuracy but is limited to short fragments, complicating the assembly of repetitive genomic regions.
  2. PacBio's HiFi technology loops DNA to read it multiple times, achieving short-read accuracy over fragments up to 25,000 base pairs long.
  3. Oxford Nanopore measures electrical disruptions as DNA passes through a pore, enabling ultra-long reads and direct detection of epigenetic modifications.
  4. No single platform captures all genomic data perfectly; clinical and research applications increasingly rely on combining data from multiple architectures.
150–300 bp
Standard Illumina read length
10,000–25,000 bp
PacBio HiFi read length
>99.9%
Accuracy of synthesis and HiFi methods
2+ million bp
Maximum continuous nanopore read

How we got here

  1. 2019

    PacBio introduces HiFi sequencing, utilizing circular consensus to achieve 99.9% accuracy on long reads.

  2. 2024

    Comparative studies confirm that long-read technologies are essential for assembling complex structural variants missed by short-read platforms.

Sources

Source coverage

8 outlets

3 viewpoints surfaced

Short-Read Advocates 40%Long-Read Specialists 35%Direct-Detection Pioneers 25%
  1. [1]IlluminaShort-Read Advocates

    Next-Generation Sequencing (NGS)

    Read on Illumina
  2. [2]PacBioLong-Read Specialists

    How HiFi sequencing works

    Read on PacBio
  3. [3]PMC

    Next-Generation Sequencing (NGS): Platforms and Applications

    Read on PMC
  4. [4]Technology NetworksDirect-Detection Pioneers

    Oxford Nanopore Sequencing: Principles, Performance, and Applications

    Read on Technology Networks
  5. [5]BiocompareShort-Read Advocates

    Next-Generation Sequencing Technology Update

    Read on Biocompare
  6. [6]MDPIDirect-Detection Pioneers

    Evaluation of Illumina and Oxford Nanopore Sequencing for the Study of DNA Methylation in Alzheimer's Disease and Frontotemporal Dementia

    Read on MDPI
  7. [7]PacBioLong-Read Specialists

    Sequencing 101: Comparing long-read sequencing technologies

    Read on PacBio
  8. [8]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Science stories with full source coverage and perspective breakdowns delivered to your inbox.