Biohub Leads $1.8B Coalition to Build Predictive AI Virtual Cells
Biohub, federal agencies, and tech leaders have pooled $1.8 billion in compute, data, and lab funding to train AI models that simulate cellular biology.

Key takeaways
- Biohub, the US Department of Energy, and the NIH are coordinating a $1.8 billion initiative to generate AI-ready biological data.
- Meta, Google DeepMind, and Isomorphic Labs are collectively contributing $300 million to the effort.
- Commercial backers receive a one-year embargo period of exclusive access before their sponsored datasets become open to the public.
- The initiative aims to scale biological training data from hundreds of millions of cells to billions and trillions over five years.
Biohub, the nonprofit biomedical research organization co-founded by Mark Zuckerberg and Priscilla Chan, announced on October 7, 2026, a massive expansion of its effort to build predictive AI models of cellular behavior. Backed by federal agencies and top technology firms, the collective commitment reaches $1.8 billion across direct funding, computing infrastructure, lab measurements, and biological repositories.
The effort unites Biohub with the US Department of Energy (DOE), the National Institutes of Health (NIH), Meta, Google DeepMind, and AI drug discovery startup Isomorphic Labs, according to an announcement from Biohub. By generating standardized, multimodal biological datasets at scale, the partners aim to construct a "virtual cell" that enables scientists to simulate cellular responses and run biological experiments computationally rather than entirely in wet labs.

Funding Breakdown and Institutional Roles
The $1.8 billion commitment pulls together private technology investments, philanthropic backing, and federal resources, as reported by TNW. Biohub anchors the initiative with $500 million pledged in April 2026 under its Virtual Biology Initiative. Of that founding sum, $400 million funds advanced cellular measurement technologies—such as cryo-electron tomography and high-throughput tissue microscopy—while $100 million supports outside academic research.
Federal agencies represent a major portion of the resource pool. The Department of Energy is committing more than $500 million over five years through its Genesis Mission. As detailed by Biohub, the DOE will deploy exascale supercomputing, autonomous laboratories, cryo-electron microscopy, and beamline facilities across its National Laboratory network to generate measurements and run biological simulations.
The NIH is contributing datasets and knowledge bases created from more than $500 million in prior federal investments through its Bio Genesis Mission. Biohub plans to work directly with the agency to clean and standardize these repositories into unified formats suitable for training deep learning models.
On the private sector side, Google DeepMind, Isomorphic Labs, and Meta are collectively providing $300 million to build foundational data and model architectures. In addition, Nvidia is providing accelerated computing infrastructure, specialized software, and technical expertise, while Renaissance Philanthropy is assisting in fundraising efforts, as noted by Biohub.

Data Access Rules and the One-Year Commercial Embargo
While the long-term goal of the initiative is an open data commons for worldwide research, private contributors will receive an early advantage. Commercial funders—Meta, Google DeepMind, and Isomorphic Labs—will have exclusive access to the datasets they fund for one year before the information is released publicly, according to Biohub Head of Science Alex Rives in an interview cited by The Decoder.
Rives explained that the embargo period serves as an incentive to attract large corporate balance sheets into open-science initiatives. Data produced through US government programs will not have an embargo period and will be made available without commercial restrictions, as reported by ksl.com.
Consortium members including the Allen Institute, the Broad Institute, the Gladstone Institutes, the Wellcome Sanger Institute, and the Human Cell Atlas and Human Protein Atlas initiatives will also help establish common standards, metadata identifiers, and unified access points for the resulting libraries.
Closing the Biological Data Gap
Current biological databases contain measurements spanning hundreds of millions of individual cells. However, Rives stated to ksl.com that accurate predictive models will require training on billions and eventually trillions of cells. Much of this data will be gathered using spatial transcriptomics, which maps molecular activity inside intact tissues, along with high-throughput perturbation screens that document how cells respond to drugs or genetic alterations.
Biohub aims to release its first coordinated dataset within approximately one year, targeting high-accuracy predictive virtual cell models within five years. According to Rives, compressing this work into a five-year window requires coordinated measurement pipelines that no single academic lab could sustain on its own.
The Race to Build the AI Virtual Cell
The initiative comes as major technology companies and research organizations expand their investments in biological foundation models. As reported by The Decoder, Anthropic has set up its own wet lab to generate biological data for drug discovery, while the OpenAI Foundation launched a grant fund exceeding $125 million to support biological and medical data collection. In Europe, Paris-based startup Rivercell announced a $25 million funding round on October 7, 2026, to develop an AI virtual cell platform.
Biohub is attempting to coordinate these disparate efforts under a unified open data architecture. If the five-year roadmap succeeds, researchers could conduct simulated biological interventions digitally to identify drug targets and evaluate cellular responses before stepping into physical laboratories, as noted by aa.com.tr. Biohub plans to approach commercial pharmaceutical companies and additional philanthropic foundations next to further expand the data generation pipeline.
Frequently asked questions
What is the goal of Biohub's Virtual Biology Initiative?
The initiative aims to generate standardized, large-scale biological datasets to train AI models that can accurately predict cellular behavior and simulate experiments digitally.
Who is funding the $1.8 billion project?
Funding includes $500 million from Biohub, more than $500 million over five years from the US Department of Energy, over $500 million in prior dataset investments from the NIH, and $300 million combined from Meta, Google DeepMind, and Isomorphic Labs.
Will the virtual cell data be made public?
Yes. Government-funded data will be accessible openly without restrictions, while commercial funders receive a one-year exclusive access period to the data they finance before it is released to the public.
When are the first datasets and models expected?
Biohub plans to release its first standardized dataset in about one year, with target predictive cell models expected within five years.
Sources
- Biohub, Meta, Google DeepMind and the US pool $1.8bn for AI biology dataTNW | Artificial-intelligence · Oct 7, 2026
- Google invests millions in Mark Zuckerberg’s efforts to create a ‘virtual cell’The Verge · Oct 7, 2026
- Zuckerberg's Biohub leads a $1.8 billion push to build AI models that predict cell behaviorThe Decoder · Oct 7, 2026
- Open data for predictive AI models of biology: $1.8 billion committedBiohub · Oct 7, 2026
- US government, Google join Zuckerberg-backed Biohub in $1.8 billion push for AI biology dataksl.com · Oct 7, 2026
- Zuckerberg's Biohub expands virtual cell effort to $1.8B, with Meta, GoogleEndpoints News · Oct 7, 2026
- Zuckerberg's Biohub teams with Google, US government in $1.8B push to model human cells with AIaa.com.tr · Oct 7, 2026
How this story was made: the newsroom picked it up from Techmeme, theverge.com and Reddit, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (42 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 8, 2026 at 01:08 UTC


