The news
Biohub, the nonprofit research group backed by Mark Zuckerberg, said on October 7, 2026 that the US Department of Energy, the National Institutes of Health and new corporate partners are joining its AI biology data initiative, bringing total investment to $1.8 billion. The goal is to generate data to train AI models that predict how cells behave.
The Department of Energy will put in more than $500 million over five years, including microscopy equipment, according to SiliconANGLE. Google DeepMind, drug discovery company Isomorphic Labs and Meta Platforms (META) are contributing $300 million combined. The NIH is contributing access to existing datasets built with more than $500 million in earlier federal funding.
Reports differ on Biohub's own share. The Next Web said Biohub committed $500 million when it launched its Virtual Biology Initiative, with $400 million for new cell measurement tools and $100 million for outside research. SiliconANGLE described a $400 million Biohub commitment, with 75% for in-house research and 25% for external projects.
Nvidia (NVDA) is providing accelerators and specialized software, SiliconANGLE reported. The Next Web listed other partners including the Allen Institute, Broad Institute, Gladstone Institutes, Wellcome Sanger Institute, Human Cell Atlas, Human Protein Atlas and Renaissance Philanthropy.
The effort will lean on three imaging methods, according to SiliconANGLE: neutron scattering, X-ray microscopy and cryo-electron microscopy. Biohub plans to use cryo-electron tomography to capture images containing millions to billions of cells.
The Next Web reported that the first dataset is expected in about a year, with accurate predictive models targeted within five years. The data will be a public resource, but commercial funders get one year of exclusive access before release; government-funded work carries no embargo.
The numbers
- Total investment
- $1.8 billion
- Department of Energy
- More than $500 million over five years
- Google DeepMind, Isomorphic Labs, Meta
- $300 million combined
- Prior federal funding behind NIH datasets
- More than $500 million
- Commercial funders' exclusive access window
- One year
Why CEOs should care
For pharmaceutical and biotech leaders, the one-year exclusive window is the key term. Google DeepMind, Isomorphic Labs and Meta will see commercially funded data a year before competitors, though government-funded work carries no embargo, which could matter in a field where AI models for drug discovery race to train on new data. Companies not in the consortium should ask whether there is a way to join as funders, and plan to use the public releases as they arrive.
For CFOs and strategy teams, the initiative may lower the cost of early experiments. If AI models can simulate how cells respond to changes, some lab testing could move to computers, as Biohub's head of science, Alex Rives, argued in describing digital experiments. That is a five-year target, not a near-term saving, so treat it as a planning scenario rather than a budget assumption.
For technology buyers and CIOs at research organizations, expect demand for imaging data storage, compute and data standards to grow. Ask how the NIH and Biohub will format and license the datasets, so internal pipelines can take them in without heavy rework.
The bigger picture
The deal is part of a wider shift in which governments and technology companies treat training data, not just models, as strategic infrastructure. Biohub called the combined investment the largest coordinated commitment to AI-ready biology data so far, The Next Web reported.
It also deepens ties between large technology companies and public science. Meta and Google DeepMind gain early data access, the Energy Department gains industry partners for its labs, and Nvidia supplies the hardware, adding biology to the list of fields where compute demand is rising.
What’s next
Watch for the first dataset, expected in about a year, the licensing terms for public release, and whether other drug makers or cloud providers join as funders to gain the early-access window.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error









