AlphaGenome Atlas Precomputes Predictions for Nearly Every Single-Letter DNA Change
Google DeepMind’s AlphaGenome is designed to predict how individual DNA changes affect molecular processes such as gene expression and splicing, not to predict a person’s disease risk. Pushmeet Kohli, who described the model, argues that its combination of a million-base-pair context window and single-base precision can help researchers investigate variants whose effects remain poorly understood. The accompanying AlphaGenome Atlas makes predictions for roughly nine billion possible single-letter changes available for research, while linking them to whole-person outcomes remains an open task.

AlphaGenome predicts molecular effects, not what a person will become
AlphaGenome predicts how changes in DNA may affect molecular properties of a cell. That is narrower than predicting a person’s disease risk, but it addresses a basic problem in biology: genomes differ, and researchers still need to understand what many of those differences do.
Pushmeet Kohli describes the genome as a recipe containing both instructions for making proteins and instructions for regulating when, where, and in what quantity they are expressed. The human genome has about three billion characters, and the human proteome includes roughly 20,000 proteins. The coding parts of the genome specify proteins; the non-coding parts help regulate their expression. Kohli calls that regulatory territory the “dark part” of the genome because it is much less well understood. The Human Genome Project made it possible to read the recipe book for a human genome, he said; the continuing challenge is to decipher what its variations do.
AlphaGenome tries to predict the molecular consequences of changing an individual letter in the genome. Every person has a genome, and those genomes are similar but not identical. The model is intended to help researchers ask what could happen if one character changes: what effect might that have on the cell?
Some individual effects are already well understood. Kohli cited the mutation associated with sickle cell anemia. But many traits and diseases involve multiple mutations rather than a single change with a clear effect. AlphaGenome focuses especially on the molecular consequences of individual variants. Connecting those predictions to the behavior of a whole organism—to disease susceptibility or other human characteristics—still requires further research, Kohli said.
That distinction matters when asking why people respond differently to similar circumstances. Károly Zsolnai-Fehér described studies in which people of similar age and build follow the same diet and training, yet show markedly different changes in muscle: some gain a great deal, some gain little, and one might even lose muscle. Kohli said genomic variation is part of what researchers are trying to understand. AlphaGenome’s predictions concern effects at the molecular level; how those effects influence a whole person remains a further research question.
Kohli said AlphaGenome is available to the scientific community for academic use, with model weights that researchers can use and fine-tune. He also said it was trained on both human and mouse genomes. In his account, training across species can make the model less likely to memorize mechanisms that apply only to human data. Instead, it must account for effects in both genomes, which may encourage it to learn broader principles. He described this as transfer learning: experience with one problem can help with a related one.
A million-base-pair view meets single-base precision
A variant’s effects may depend on DNA beyond its immediate neighborhood. Pushmeet Kohli said that three-dimensional interactions can make the relevant region large, so analyzing a change may require looking across a broad stretch of the genome. A model that examines only nearby bases could miss effects associated with a wider genomic context; one that considers a wide region at coarse resolution could miss the detail of an individual base change.
Earlier models such as Enformer used both a shorter context window and a coarser resolution, Kohli said. They did not examine a million-base-pair window, and they grouped base pairs rather than making predictions at the level of each individual base. AlphaGenome’s advance, as he described it, is to combine a one-million-base-pair context window with base-pair resolution. The broad window gives the model a larger neighborhood to consider, while the fine resolution lets it make predictions about individual positions within that neighborhood. The two capabilities address different constraints: how much surrounding sequence is available, and how precisely the sequence is represented.
Károly Zsolnai-Fehér cited the paper’s comparison against other models: AlphaGenome matched or outperformed the strongest comparison model on 25 of 26 variant-effect benchmarks. He also cited a reported correlation of about 0.5 between predicted and observed effect sizes for mutations affecting gene activity. Kohli agreed that the model can predict not only whether a mutation matters but, in some cases, the rough magnitude of its effect.
Kohli described the AlphaGenome Variant Impact score, or AVI score, as a way to help scientists explore the large amount of predicted data. A high score, he said, identifies a variant that researchers should consider and analyze carefully. A low score, in his description, indicates high confidence that the variant will not have a dramatic impact on the phenotypes assessed. The score is a prioritization aid for examining predicted effects.
The team’s confidence in the evaluation results came after years of work, Kohli said. The project had been under way for four or five years or more, involving work on data, architecture, and training. When results began to improve, the team checked that the metrics and evaluations were valid. Kohli said the team had confidence in those measures partly because they had been developed over the course of the project, rather than only after the results appeared.
The Atlas turns variant predictions into a resource for researchers
Pushmeet Kohli described AlphaGenome as answering a targeted question: what might happen if a particular base pair changes—for example, from A to T, or from T to C? The Atlas applies the model at a far larger scale. Kohli called it a dictionary of what happens when variation occurs in the genome.
There are roughly three billion positions in the human genome, and each position can be changed to one of three other letters. That yields about nine billion possible single-nucleotide variations. For each, AlphaGenome can produce predictions about molecular effects, including effects on splicing and gene expression. The Atlas precomputes those predictions and makes the resulting dataset available to the scientific community.
That precomputation changes what a researcher can do with the model. Instead of posing a single variant question and waiting for a prediction, researchers can explore a large, organized body of predictions across possible substitutions. Kohli described the result as a petabyte-scale dataset, containing detailed information about the effects of variation. The Atlas also includes the AlphaGenome Variant Impact score, which helps researchers identify predictions that merit closer analysis.
The scale recalls the AlphaFold database, which made protein-structure predictions available at scale. Kohli said AlphaFold’s database contained predictions for 250 million proteins, representing almost all the proteins known at the time, and had been used by more than four million people across 190 countries. He said AlphaGenome Atlas is 30 times the size of the AlphaFold database.
For Kohli, the Atlas belongs to a class of foundational resources that can support work in many directions. He called deciphering the genome a “root node” problem: a capability whose implications extend beyond a single application. Understanding disease mechanisms and identifying variations that control the expression of proteins in a disease pathway are among the possibilities he described. The model could also help researchers study genomes or proteins developed for synthetic biology. In that setting, he said, AlphaGenome offers a way to examine the properties of those genomes, including genomes sampled using DNA language models such as Evo.
Molecular predictions need biological context and experimental testing
Károly Zsolnai-Fehér asked how researchers could distinguish a model that has learned biological mechanisms from one that has learned statistical correlations. Pushmeet Kohli agreed that this is a central challenge in computational biology. His answer focused on what AlphaGenome predicts: molecular properties such as splice sites, splicing, and gene expression, rather than a direct relationship between a DNA change and a disease diagnosis. Kohli said these molecular-level predictions can be explained in biophysical terms.
He also pointed to laboratory validation. When researchers take predictions and test them in the lab, Kohli said, the results make sense and can lead to novel findings about the importance of particular variants. He said the tasks are grounded in physics, while how the model itself reaches its predictions remains an area the team is trying to interpret.
The model’s simple final prediction layer prompted a related exchange. Zsolnai-Fehér noted that the paper uses Lasso regularization on top of the more complex model, likening it to a small bicycle bell attached to a Formula One engine. Kohli said the simplicity illustrates the value of the model’s learned representation: most of the work should happen in the way AlphaGenome represents the DNA sequence, allowing the predictor on top to be relatively simple.
The next steps Kohli described address both the model’s predictions and their biological context. Different cells behave differently, so the team is working to make AlphaGenome more aware of the context in which it makes predictions. Kohli also cited expanding the training data to improve accuracy across different properties, and helping researchers connect the model’s molecular phenotypes to organism-level outcomes such as disease propensity.
The larger ambition is to make complex biology more tractable
Asked what further progress in AI might make possible, Pushmeet Kohli pointed to simulating a cell—a “virtual cell”—and eventually understanding or simulating an organism. He said a deeper understanding of cellular workings could have consequences for cancer research and synthetic biology. These are longer-range scientific ambitions, not capabilities he attributed to AlphaGenome today.
Kohli also named understanding long-term climate effects and developing materials that store or transform energy more efficiently as scientific problems that greater capability might help address. He framed these alongside the challenge of understanding biology from first principles and designing biological functions relevant to human health and agriculture. The examples extend beyond genome prediction, but share the broader aim of making complicated systems more understandable and useful.
The distinction between a useful demonstration and a complete account came up in Kohli’s description of Gemini. He said he had asked it to explain how engineers use wind tunnels to study airplane parts. The system created a simulation that let his son adjust parameters and see changes in airflow. Kohli said it was not completely accurate, but considered it a useful way to show a theory in action.
The broader ambitions are set against scientific questions that remain unresolved. Kohli cited understanding the brain and explaining material properties. Even for superconductivity, he said, there is not a complete theory. In this framing, AI’s potential is to help make difficult systems more tractable to study.

