E. coli RNA Polymerase Successfully Transcribes Synthetic Eight-Letter DNA Alphabet
High-resolution imaging has proven that a natural bacterial enzyme can accurately read and transcribe an expanded genetic code, including synthetic base pairs that lack hydrogen bonds.
- Synthetic Biologists
- View the expanded alphabet as a programmable toolkit to engineer cells that produce novel therapeutics, materials, and proteins beyond the limits of natural evolution.
- Structural Biologists
- Focus on the mechanistic revelation that enzyme catalysis relies more on precise spatial geometry and trigger-loop closure than on specific hydrogen-bonding chemistry.
- Evolutionary Theorists
- Argue that the success of the eight-letter system proves the four-letter code of Earth's life is an evolutionary accident rather than a strict chemical necessity.
Perspectives this story doesn't cover
- Bioethicists
- Regulatory Agencies
- 8 letters
- Size of the expanded Hachimoji DNA alphabet
- 2.42–2.75 Å
- Resolution of the cryo-EM structures capturing the enzyme
- 2x slower
- Transcription speed of the synthetic P:Z pair compared to natural G:C
- 3.20 Å
- Resolution of the structure showing hydrophobic base pair transcription
For decades, a central assumption in molecular biology held that the chemistry of life was strictly constrained by the four natural DNA bases—adenine (A), thymine (T), cytosine (C), and guanine (G)—and the specific hydrogen bonds that hold them together. The opposing view, championed by synthetic biologists, argued that the genetic code is merely a chemical scaffold. If synthetic molecules could mimic the geometry of natural base pairs, they hypothesized, the cellular machinery would process them just the same, regardless of their chemical makeup.[4]
Now, high-resolution structural evidence has settled the debate. In a pair of studies published in late summer 2026, researchers at the University of California San Diego demonstrated that Escherichia coli RNA polymerase—the enzyme responsible for reading DNA and synthesizing RNA—can accurately transcribe an expanded, eight-letter genetic alphabet.[1][2]
The expanded system, known as Hachimoji DNA (from the Japanese words for "eight" and "letters"), doubles the natural genetic code by introducing four synthetic nucleotides: P, Z, B, and S.[1]
In a study published September 2 in Nature Communications, the research team, led by Dong Wang, used cryo-electron microscopy to capture the bacterial enzyme in the act of transcribing the synthetic base pairs.[2]
The structural snapshots, resolved to between 2.42 and 2.75 angstroms, revealed that RNA polymerase recognizes the synthetic letters using the exact same biochemical and structural signals it uses for natural DNA. The enzyme's active site adopts a catalytically competent configuration, seamlessly incorporating the artificial bases into the growing RNA strand.[2]
The enzyme's active site adopts a catalytically competent configuration, seamlessly incorporating the artificial bases into the growing RNA strand.
"By expanding the genetic code, we could create new molecules that have never been seen before and explore new ways of making proteins as therapeutics," Wang stated.[1]
The flexibility of the transcription machinery extends even further than geometric mimicry. In a companion study published August 12 in the Proceedings of the National Academy of Sciences, the same team showed that RNA polymerase can process a hydrophobic unnatural base pair—known as Ds:Pa—that completely lacks the hydrogen bonds normally required to hold DNA strands together.[3]
The PNAS study captured the enzyme at a 3.20-angstrom resolution, showing that the hydrophobic Ds:Pa pair forms an edge-to-edge alignment. The enzyme's "trigger loop"—a critical structural element for catalysis—closes fully around the synthetic pair, proving that hydrogen bonding is not a strict prerequisite for transcription.[3]
However, the process is not entirely symmetrical. The kinetic data revealed that the enzyme incorporates the synthetic Ds nucleotide much more efficiently than its partner, Pa. Similarly, the Nature Communications study noted that the synthetic P:Z base pair is processed at a transcription speed roughly two times slower than a natural G:C pair.[2][3]
Despite these kinetic speed limits, the overall fidelity remains high. The findings provide the first structural proof that a parallel, alternative genetic system can be fully supported by natural cellular enzymes.[1][4]
This structural foundation moves synthetic biology closer to practical applications. By proving that living cells can process an eight-letter code, researchers can now design synthetic DNA templates that instruct cells to manufacture novel amino acids, complex diagnostics, and targeted cancer therapeutics that the natural four-letter alphabet could never encode.[1][4]
What we don’t know
- It remains unclear how the slower transcription kinetics of synthetic base pairs will affect overall cell viability and growth rates in a fully living organism.
- Researchers do not yet know the upper limit of how many synthetic base pairs a single cell's machinery can tolerate before transcription errors become fatal.
- The long-term evolutionary stability of the Hachimoji system inside a dividing, living cell has not been fully mapped.
Sources
[1]ScienceDailySynthetic BiologistsLife uses 4 DNA letters. Scientists just made 8 work
Read on ScienceDaily →
[2]Nature CommunicationsStructural BiologistsStructural basis of transcription of the hachimoji eight-letter alphabet by E. coli RNA polymerase
Read on Nature Communications →
[3]Proceedings of the National Academy of SciencesStructural BiologistsHydrophobic unnatural base pair promotes trigger loop closure and catalysis in cellular RNA polymerase independent of hydrogen bonding
Read on Proceedings of the National Academy of Sciences →
[4]Factlen Editorial TeamEvolutionary TheoristsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Science
See all →Water Quality
How the Maximum Contaminant Level Balances Health Risk and Economic Feasibility in Drinking Water
8 sources
Population Genetics
Calculating the Hidden Carriers: How the Hardy-Weinberg Equation Maps Population Genetics
6 sources
Island Biogeography
Island Size and Distance: How the Equilibrium Model of Biogeography Predicts Species Richness
8 sources
Cellular Biology
How the Human Body Replaces 330 Billion Cells Every 24 Hours
6 sources
Every angle. Every day.
Get Science stories with full source coverage and perspective breakdowns delivered to your inbox.




