The human genome contains about three billion letters. For each position, there are potential substitutions that, considered on a genomic scale, yield approximately nine billion single-nucleotide variants. Testing them one by one in the laboratory is simply impossible. Google DeepMind decided to tackle the problem in what has become the typical approach of scientific AI: precomputing a prediction for all of them.
On September 8, DeepMind unveiled AlphaGenome Atlas, a platform containing predictions on the molecular effects of roughly 9 billion possible single-letter changes in human DNA. The dataset occupies about one petabyte and is, according to the company, more than thirty times larger than the AlphaFold Database. It is freely available for academic research via a web portal and an API.
The news is not merely that a model can generate an enormous number of predictions. The point is that DeepMind is attempting to turn a complex model into scientific infrastructure. It is the same dynamic seen with AlphaFold: first, demonstrate that AI can predict something useful; then, precompute results at massive scale and build an archive that allows thousands of researchers to use those predictions without needing the infrastructure required to generate billions of them.
From model to atlas
AlphaGenome had already been introduced as a model capable of predicting how a genetic variant might influence biological processes such as gene expression, RNA splicing, and other molecular signals. However, querying the model variant by variant remains computationally expensive and demands technical expertise.
The Atlas flips the workflow. DeepMind ran the predictions in advance and organized them into a searchable catalog. Instead of asking the model to compute the effect of a single mutation each time, researchers can look up a genomic position and immediately obtain an estimate of the biological processes that might be altered.
This is a significant shift because much of modern genetics faces a prioritization problem. Sequencing generates massive volumes of variants; the majority have no immediately clear clinical or biological significance. The challenge is not finding differences in DNA, but figuring out which ones genuinely warrant an experiment.
The AVI score aims to turn complexity into a ranking
Alongside the Atlas, DeepMind introduced the AlphaGenome Variant Impact score, or AVI. The score combines insights from AlphaGenome with those from AlphaMissense—the company’s model focused on protein-altering variants—and condenses multiple predictions into a single impact metric.
The goal is to allow researchers to quickly rank the most promising variants. The system is not limited to the small portion of the genome that directly codes for proteins: DeepMind emphasizes that it can also be used in non-coding regions, which make up about 98% of the genome and contain many regulatory elements governing gene activity.
This is precisely where such tools can prove invaluable. Variants in coding regions are often more intuitive to interpret: they alter a protein, and its effect can be studied. Regulatory regions, by contrast, are far more challenging. A mutation can alter when a gene is activated, how much it is expressed, or how its RNA is processed, without directly changing the final protein sequence.
The map is not the territory
The communication risk with AlphaGenome Atlas is the same one that accompanies much of scientific AI: mistaking a prediction for an experimental discovery.
The Atlas does not prove that a specific variant causes a disease. It predicts how that variant might influence molecular processes. Moving from a predicted correlation to biological causality requires experiments, clinical data, and individual context. In covering the launch, Nature highlights this very limitation through independent researchers: such a tool can help narrow the field, but it does not replace the wet lab or the evaluation of a clinical case.
It is a crucial distinction, especially in rare diseases. A patient may carry thousands of variants compared to a reference genome. A system that ranks those most plausibly harmful can drastically cut research time, but a diagnosis cannot be automatically outsourced to an AI-generated score.
The economic value lies in time saved
The most tangible transformation could occur long before any AI-designed therapies emerge. If a laboratory needs to select ten variants to test out of ten thousand candidates, improving the quality of that selection means saving months of work, reagents, personnel, and experimental bandwidth.
In biology, the cost is not purely computational. Every hypothesis that enters the lab competes with others for instruments, samples, and researchers' time. A better ranking system does not have to be perfect to be useful: it simply needs to increase the probability that the first hypotheses tested are the right ones.
DeepMind reports that external collaborators have already used AlphaGenome Atlas to identify and subsequently verify experimentally relevant variants in rare disease cases, as well as to study associations with common traits. Among the examples, Nature cites researchers at the Broad Institute using the system to prioritize a non-coding variant as a potential explanation for a case of severe epilepsy.
The AlphaFold effect
The inevitable precedent is AlphaFold. The most significant leap of that project was not just a model capable of predicting protein structures, but the decision to make a database containing hundreds of millions of predicted structures available at scale.
Once the cost of accessing predictions collapses, AI stops being a laboratory experiment and becomes a scientific utility. DeepMind seems intent on doing the same with regulatory genetics: turning a sophisticated capability into an infrastructure queryable even by those without a machine learning team.
AlphaGenome Atlas is even more ambitious in one respect: the dataset is enormous and describes not already catalogued biological objects, but a universe of possible variations. In a sense, it is a map of experiments that have never been conducted.
The problem of trust in scientific models
The more widely used these maps become, the greater the need to understand their errors. A model can be highly accurate on average and yet fail on a rare class of variants. It can reflect biases in training data or perform differently across tissues, populations, and types of genetic regulation.
This is why a predictive atlas should be interpreted as a research support system, not as a catalogue of biological truths. Scientific value increases when predictions are continuously benchmarked against new experiments and when instances where the model errs are made visible.
Accessibility itself can have a positive effect: thousands of teams querying the Atlas will inevitably generate a massive volume of independent validations and refutations. If this data feeds back into the development cycle, the system could improve much like a major shared infrastructure.
Genomics is becoming a navigation problem
For years, the primary bottleneck was obtaining genetic data. Today, sequencing is far cheaper, and the challenge has shifted toward interpretation. We have ever more letters of DNA to read, and ever less relative capacity to understand which ones truly matter.
AlphaGenome Atlas was created precisely in this space. It does not automatically discover the cause of a disease, nor does it make geneticists redundant. It attempts to make a space of variants too vast to explore directly navigable.
If it performs well, its significance will not be measured by the number of predictions generated, but by the number of experiments it helps avoid and by how many correct hypotheses it brings to the lab bench sooner.
It is a less spectacular form of artificial intelligence than a chatbot that talks like a human. But it could prove to be one of the most profound: turning billions of biological possibilities into a map useful enough to decide where to look.



