On October 7, Biohub (the research nonprofit started by Priscilla Chan and Mark Zuckerberg), the US Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs and Meta committed $1.8 billion to build open data for AI models of the cell. They want to measure how cells behave at a scale nobody has tried, and then make the results public.
I think it's awesome. So did a lot of people: one post about it passed 1.7 million views in a day, and researchers who build these models, like Bo Wang, called it significant.
The size of the gap explains the size of the check. Today's best cell datasets hold hundreds of millions of cells, and Biohub's head of science told Reuters that a model you could trust to predict cell behavior will need billions, then trillions. You don't close that gap with one lab, or ten.
Hundreds of millions sounds like a lot until someone says trillions.
Look at where Biohub's own money goes, though. Little of it buys AI. Most of it buys microscopes and experiments.
The library behind AlphaFold
AlphaFold predicts the 3D shape of a protein from its sequence, a problem researchers had chipped at since the 1970s. It won Demis Hassabis and John Jumper the 2024 Nobel Prize in Chemistry, and more than two million people in 190 countries have used it. But the model came last. It was trained on databases of every known protein structure and sequence, most famously the Protein Data Bank, which structural biologists filled one painstaking experiment at a time over decades and gave away for free.
Nobody won a Nobel for filling in the Protein Data Bank. The people who trained a model on it did.
A field spends years building a shared library nobody gets rich from, and then a model shows up and makes the library look like the smartest investment anyone ever made.
Cells don't have that library. Biohub's Alex Rives named the problem when he launched this effort in April:
"The cell is orders of magnitude more complex, and we will need to create the data in just a few years rather than decades."
This announcement is a field trying to build its Protein Data Bank on purpose, in five years instead of fifty. The Human Genome Project is the closest precedent, and it's a good one: a deliberate, coordinated, public library that a whole industry has been building on since.
One of the companies writing the check is Google DeepMind. The lab whose model learned from the last great open science library is now paying for the next one. That's a nice way to say thank you.
What a virtual cell would do
The thing everyone here wants to build is called a virtual cell: an AI model that predicts how a cell responds when you change something. Add a drug, switch off a gene, change what's around it, and the model tells you what happens before anyone picks up a pipette.
Training one takes a mountain of examples of exactly that. Push a cell, record what changes, and repeat across enough cell types, drugs and genes that the model learns the rules well enough to predict pushes nobody has tried.
The payoff is easiest to see in the cells you can't easily experiment on. Rare cell types, cells that won't grow outside the body and diseased tissue that can't be cultured faithfully are often the ones closest to the disease you care about. A model that learned from the experiments we can run, and transfers that to the ones we can't, would open up a lot of biology that's currently off limits.
Nobody should expect this to replace the lab. It changes which experiments get run. You test a thousand ideas on the computer, and the lab gets the handful most likely to work.
Every partner brings the thing only it has
Nobody here is just writing a check. Each partner brings something the others can't.
Biohub committed $500 million in April, and $400 million of it goes to generating data and building instruments: microscopes that can image millions to billions of cells in living tissue, imaging that resolves the inside of a cell in near-atomic detail, and tools to engineer cells for experiments. The remaining $100 million funds outside labs. Biohub has run this play before, with open cell atlases and data portals the field already uses.
The Department of Energy brings more than $500 million over five years and, more usefully, the national labs: supercomputers, imaging facilities and robotic labs that run experiments with little human help.
The NIH brings data the public already paid for, built with more than $500 million in earlier research funding, which Biohub will help standardize so models can train on it. That's the cheapest win in the whole package, and the easiest one to cheer.
Google DeepMind, Isomorphic Labs and Meta put in $300 million between them. Google and Meta splitting the bill on something neither of them gets to own isn't something you see every week. Add that to the DOE's commitment and about $800 million of the headline is new this week. The rest is Biohub's April money and existing NIH data, and I think counting both is fair.
The research institutes are in too. The Allen Institute, Broad, Gladstone, Wellcome Sanger, the Human Cell Atlas and the Human Protein Atlas are coordinating, with NVIDIA on compute. Broad and Sanger were both major contributors to the Human Genome Project, so veterans of the last deliberate public library are in the room for this one.
The commercial money comes with a head start. Rives told Reuters the corporate funders get an embargo period to work with the data before it goes public, and the government-funded work carries no such restriction. He didn't say how long the window runs. It still seems like a sensible trade to me. It's the price of getting companies to fund data that ends up public, and Biohub, Arc and Tahoe set up their January dataset the same way: shared among the three first, open-sourced after.
The small version already worked
The field has already run a small version of this, and it worked.
Tahoe Therapeutics released an open dataset of cells responding to drugs in 2025. It's been downloaded more than 250,000 times and, alongside two other open datasets, has served as training or reference data for the leading virtual cell models. In January, Tahoe, Arc Institute and Biohub announced a successor with more than 120 million cells and 225,000 drug and patient combinations, about four times richer in experiments.
The field also has a scoreboard. Arc's Virtual Cell Challenge drew more than 5,000 registrants from 114 countries in its first year, with over 1,200 teams submitting. This year's round asks models to predict six cell lines they've never seen tested, which is much closer to what a biologist would actually want. Arc calls the target "the AlphaFold and ImageNet moment" for cell modeling, then adds that it "would be pleasantly surprised if anyone reaches this level in 2026."
You don't often see an organization launch a contest and say in the same breath that the real target is probably out of reach this year. I enjoyed that, and it makes me trust the scoreboard more.
So the ingredients exist: open datasets people actually download, a public benchmark, and consortia like the Human Cell Atlas with thousands of members in over a hundred countries. What they didn't have was money and coordination at this scale. Rives told Reuters the partners want to do decades of work in five years, with the first dataset in about a year.
I have some skin in this one. In April I wrote The Data Paradox, asking whether the decade we spent cleaning data was the last decade it mattered. Models had gotten good enough at chewing through messy data that a lot of cleanup work looked less essential than it used to.
I still think that's half right, which is a generous grade to give myself. This announcement shows me the half I underplayed.
A model can work around a messy spreadsheet. It can't work around an experiment nobody ran.
You can't clean, harmonize or prompt your way to a measurement that doesn't exist. Standardizing the NIH's old data fixes the first problem. Only new experiments fix the second. The people closest to this problem looked at it and put $400 million of Biohub's first $500 million into instruments and data generation. That tells you where AI biology is stuck.
After 25 years walking R&D labs, I find that encouraging. The labs were never short of smart people with good questions. They were short of the time and money to run every experiment worth running. A good virtual cell won't replace those experiments. It'll help scientists pick which ones to run first.
The same Reuters report notes the OpenAI Foundation's grant program of more than $125 million for biological datasets, and Anthropic adding a wet lab. More of the frontier labs are putting money into new biology data.
What should we measure first?
The open question is the one Bo Wang, whose team built the single-cell model scGPT, raised on announcement day: what should biology scale? More cells, more kinds of experiments, more kinds of tissue? Language models had scaling laws that told researchers what to buy next. Biology doesn't have its own yet, so this dataset is also the experiment that finds them.
When the first batch lands, about a year from now, we'll start learning which measurements make these models smarter, and that answer will shape where the next billion goes.



