A collaborative project backed by Meta CEO Mark Zuckerberg’s nonprofit Biohub, the US government, and tech giants Google and Meta has been launched with USD 1.8 billion in funding to assemble large-scale biological datasets for artificial intelligence research. Unveiled on Wednesday, October 7, 2026, the project seeks to produce comprehensive data on cellular behavior across various conditions to aid scientists in building predictive AI models and potentially accelerating pharmaceutical development.
Named the Virtual Biology Initiative, the collaboration unites Meta, Google DeepMind, drug-discovery firm Isomorphic Labs, the US Department of Energy, and the National Institutes of Health. Biohub itself was established with financial support from Meta leader Mark Zuckerberg and his wife, Priscilla Chan.
New Datasets Could Speed Up Drug Discovery
Commitments to the project feature a five-year pledge of over USD 500 million from the Department of Energy directed toward laboratory measurements, modeling, and computing. Meta, Google DeepMind, and Isomorphic Labs are contributing a combined USD 300 million.
Furthermore, the NIH will coordinate and standardize repositories and datasets generated through more than USD 500 million in prior federal grants to prepare them for AI training. Earlier in the year, Biohub pledged an additional USD 500 million to the undertaking.
This unified push is designed to tackle a fundamental bottleneck in applying artificial intelligence to the biological sciences: the shortage of structured, high-quality biological data compared to the massive volumes utilized in training contemporary AI models.
Project Aims to Build a Virtual Model of Cells
Biohub intends to gather data detailing cellular responses to diverse biological and environmental shifts. Investigators will employ methods such as spatial transcriptomics—which details molecular activity within intact tissues—alongside large-scale assays measuring cellular reactions.
The ultimate objective is the development of predictive models capable of simulating cellular functions. Such tools could allow researchers to run computational tests on hypotheses prior to investing time and resources into physical laboratory work.
According to Biohub head of science Alex Rives, current datasets encompass hundreds of millions of cells, whereas reliable predictive systems will likely necessitate billions or trillions of data points.
Also Read: How Mark Zuckerberg is Building the Future of AI-Powered Social Media
Data Will Eventually Become Public
While structured as an open-science venture, commercial collaborators will be granted temporary early access to the datasets they co-finance. Following these embargo windows, the information will be published as an open scientific asset, whereas government-financed work will be free of such limitations.
Biohub anticipates releasing the initial dataset in approximately one year, with plans to produce progressively refined predictive models over a five-year span. This rollout coincides with heavy investments in biological data and drug discovery by other artificial intelligence enterprises and research groups. Anthropic has broadened its focus on biology, and the OpenAI Foundation has introduced a grant initiative dedicated to medical and biological datasets.




