Skip to main content
AI BiologyChan Zuckerberg Biohub· 5 min read· in Science

US Agencies and Tech Giants Commit $1.8 Billion to Build a Predictive Virtual Cell

A $1.8 billion public-private partnership aims to map the exact molecular mechanics of living tissue to train next-generation artificial intelligence. The initiative will fund advanced measurement technologies and exascale computing to help researchers test experimental treatments entirely in software.

By Mateo Ramos

To build an artificial intelligence model capable of predicting how a human cell reacts to a virus or a drug, researchers must map biological interactions at a scale that dwarfs the text data used to train large language models. On Wednesday, a coalition of US federal agencies, tech companies, and the Chan Zuckerberg Biohub committed $1.8 billion to generate that foundational data from scratch.[1][2][5]

The Virtual Biology Initiative aims to compress decades of laboratory trial-and-error into a digital "virtual cell" that scientists can experiment on in software. By pooling public and private resources, the project seeks to overcome the primary bottleneck in biological AI: the sheer lack of high-quality, standardized experimental data required to train predictive models.[2][5]

Unlike language models that learn from vast amounts of text already published on the internet, AI biology depends on physical measurements that often have to be created in a laboratory. Current cellular datasets hold hundreds of millions of cells, but an accurate predictive model will eventually require billions or trillions of data points.[2]

To reach that scale, the initiative will systematically expose cells to various perturbations—such as turning specific genes on and off, or introducing chemical compounds—and record the exact molecular responses. This systematic mapping will provide the foundational training data needed to teach algorithms how living systems operate under stress.[1][2]

Federal and philanthropic muscle

The $1.8 billion commitment represents one of the largest coordinated investments in open biological data to date. Biohub, the nonprofit biomedical research organization founded by Mark Zuckerberg and Priscilla Chan, anchored the effort with a $500 million pledge made in April 2026.[1][4]

The $1.8 billion initiative combines fresh capital with legacy federal investments to fund data generation and computing infrastructure.

The US Department of Energy is adding more than $500 million over five years through its Genesis Mission. This federal contribution will provide the initiative with access to exascale supercomputers, X-ray and neutron scattering facilities, and autonomous laboratories across the US National Laboratory system.[1][4][5]

The National Institutes of Health is not writing a fresh check, but is instead pooling datasets and repositories built with more than $500 million in prior federal funding. Biohub will work directly with the agency to standardize those legacy resources so they can be ingested by modern AI training algorithms.[1][4][5]

Standardizing legacy data is a notoriously difficult challenge in computational biology, as different laboratories historically used varying formats and measurement standards. By unifying these disparate repositories into a single access layer, the NIH and Biohub hope to create a cohesive foundation for the new predictive models.[4]

The technology sector's role

To bridge the remaining funding and computational gaps, major technology firms are stepping into the biological domain. Meta, Google DeepMind, and Alphabet's drug discovery spinout Isomorphic Labs are collectively investing $300 million to develop multimodal datasets and predictive models alongside the public agencies.[1][2][4]

Nvidia is also joining the coalition, contributing accelerated computing infrastructure, software, and technical expertise to handle the massive data loads. The involvement of these tech giants underscores a broader industry shift, as companies that built their valuations on consumer software increasingly view biology as the next frontier for artificial intelligence.[1][4]

Biohub is directing $400 million of its own share toward advanced measurement tools to feed these systems. This includes deploying cryo-electron tomography to resolve near-atomic details inside cells, and utilizing high-throughput microscopy capable of observing millions of cells simultaneously in living tissue.[1][4]

By systematically recording how cells respond to chemical and genetic changes, researchers aim to teach AI algorithms the underlying rules of cellular biology.

These advanced imaging techniques will allow researchers to capture the spatial organization of molecules within a cell, rather than just cataloging their presence. Understanding where proteins are located and how they interact in three-dimensional space is critical for building a virtual model that accurately reflects physical reality.[4]

The commercial compromise

While the initiative's ultimate goal is to create an open public resource for the global scientific community, the data will not be immediately free to everyone. The commercial funders, including Meta and Google DeepMind, will receive a one-year period of exclusive access to the datasets they help generate.[1][2][5]

Biohub Head of Science Alex Rives defended the embargo period as a necessary compromise to secure private-sector funding. Rives noted that the exclusivity window creates a tangible incentive for commercial players to participate, ensuring the project reaches the scale required to succeed.[1][2]

After the one-year head start expires, the datasets will be released publicly to academic researchers and rival pharmaceutical companies. Government-funded work generated through the Department of Energy and the National Institutes of Health carries no such commercial restrictions and will be integrated directly into the open repositories.[1][2]

This hybrid funding model reflects the immense cost of modern biological research, where neither philanthropic organizations nor federal agencies can easily shoulder the entire burden alone. By granting a temporary commercial advantage, Biohub aims to accelerate the overall timeline for public data release.[2]

Illustration: The US Department of Energy is providing access to exascale supercomputers to process the massive biological datasets generated by the initiative.

Simulating disease digitally

If successful, the virtual cell could fundamentally reshape how pharmaceutical companies and academic researchers approach drug discovery. Instead of spending years testing compounds on physical cell cultures in a laboratory, scientists could simulate those interactions digitally to predict which interventions will work.[2][5]

"An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally," Rives said in a statement. He described the creation of a virtual cell as one of the most important challenges for the next era of science.[4][5]

The coalition expects to release its first major dataset within a year, targeting the completion of accurate predictive models within five years. The project will also draw on expertise from major independent research institutes to ensure the models reflect a comprehensive understanding of human biology.[1][2][5]

The transition from physical pipettes to digital simulations now depends on how faithfully these new algorithms can reproduce the chaotic complexity of living systems. If the models prove reliable, they could drastically reduce the failure rate of experimental drugs before a single compound ever reaches human clinical trials.[2][5]

Key points

  • A coalition of US agencies, tech giants, and Biohub committed $1.8 billion to generate AI-ready biological data.
  • The initiative aims to build a "virtual cell" that allows scientists to simulate how human tissue responds to experimental drugs.
  • The funding covers advanced measurement tools, exascale supercomputing, and the standardization of legacy federal datasets.
  • Commercial partners will receive a one-year period of exclusive access to the new data before it is released publicly.

Open questions

  • Whether AI models trained on isolated cellular data can accurately predict complex, whole-organism responses in human patients.
  • How the one-year commercial embargo will affect the pace of academic research relying on the new datasets.
  • Which specific diseases or drug classes the initial predictive models will prioritize for simulation.
Open Science Advocates 35%Commercial AI Developers 35%Federal Research Agencies 30%
Open Science Advocates
Argue that foundational biological data should be immediately accessible to the global research community.
Commercial AI Developers
Maintain that temporary data exclusivity is necessary to justify the massive capital investments required.
Federal Research Agencies
Focus on leveraging national infrastructure to maintain leadership in biotechnology and artificial intelligence.

Perspectives this story doesn't cover

  • Academic researchers subject to the data embargo
  • Patient advocacy groups awaiting new treatments

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Open Science Advocates 35%Commercial AI Developers 35%Federal Research Agencies 30%
  1. [1]Endpoints NewsCommercial AI Developers

    Zuckerberg's Biohub expands virtual cell effort to $1.8B, looping in Isomorphic, Meta, Google

    Read on Endpoints News →
  2. [2]AxiosCommercial AI Developers

    Zuckerberg's $1.8B AI project seeking to build a virtual cell

    Read on Axios →
  3. [3]QuartzFederal Research Agencies

    Google, Meta, and U.S. government join Biohub's $1.8 billion AI biology data push

    Read on Quartz →
  4. [4]Pulse 2.0Federal Research Agencies

    Biohub, DOE, NIH And Partners Commit $1.8 Billion To Build AI-Ready Biology Data

    Read on Pulse 2.0 →
  5. [5]The American BazaarOpen Science Advocates

    Biohub launches $1.8 billion AI biology initiative with Google, Meta, US government

    Read on The American Bazaar →

Comments

Stay informed

Every angle. Every day.

Get Science stories with full source coverage and perspective breakdowns, free every day.