# AlphaGenome Atlas: A Comprehensive Guide to the Genome-Wide Variant Impact Resource
## Introduction
A major resource has emerged from Google DeepMind that aims to transform how researchers approach genetic variation. AlphaGenome Atlas provides precomputed predictions for approximately 9 billion single-nucleotide changes across the human genome, effectively creating a massive lookup table for variant effects. This catalog, accompanied by the AlphaGenome Variant Impact (AVI) score, offers a new paradigm for studying how genetic mutations influence molecular processes like gene expression, RNA splicing, and protein function.
## What Makes AlphaGenome Atlas Different
Traditional approaches to studying genetic variants require running computational models on demand for each variant of interest—a slow process when analyzing genome-scale datasets. AlphaGenome Atlas solves this bottleneck by precomputing all variant effects in advance, storing results in a 1-petabyte dataset that is more than 30 times larger than the AlphaFold Database of protein structures.
The resource shifts the unit of work from “query one variant at a time” to “search across the entire genome instantly.” Researchers can now look up predictions for any single-letter DNA change and receive not just a single score, but a detailed breakdown of predicted molecular effects across hundreds of human and mouse cell types and tissues.
## The Four Core Resources
AlphaGenome Atlas packages four interconnected data layers:
**Molecular effect predictions** — Thousands of predictions per variant covering multiple aspects of gene regulation. These predictions span diverse cell types and tissues, giving researchers context-specific insights into how a variant might alter biological function.
**The AVI score** — A single impact ranking number that combines regulatory predictions from AlphaGenome with protein-altering predictions from AlphaMissense. This unified score works across the entire genome, including both the 2% that codes for proteins and the remaining 98% that controls gene regulation.
**Feature attributions** — Each AVI score is decomposed into additive contributions from interpretable biological processes such as chromatin accessibility, splicing efficiency, and evolutionary conservation. This transparency allows scientists to understand not just whether a variant is impactful, but why.
**DNA sequence motifs** — A compendium of over 2,500 recurrent short sequences found throughout the genome, complete with genomic locations and annotations for transcription factor binding sites. This resource helps researchers understand the regulatory grammar of the genome.
## How the Resource Was Built
The AlphaGenome model, first introduced in mid-2025, predicts how DNA variants change molecular processes. To create the Atlas, researchers ran this model systematically across all 9 billion possible single-nucleotide variants in the human genome and stored the resulting predictions. The precomputation approach removes the need for real-time model inference, enabling rapid queries across the entire genome.
The AVI score was designed to integrate two complementary prediction systems: one focused on regulatory effects and another focused on protein-altering changes. This integration means the score remains meaningful whether a variant falls in a coding region or a non-coding regulatory region.
## Early Research Applications
Several research teams have already leveraged AlphaGenome Atlas for novel discoveries before the resource’s public launch.
**Rare disease insights** — A collaboration with the GREGoR Consortium led to the discovery of a variant in the DNM1 gene, strongly linked to epileptic encephalopathy. The Atlas predicted that this variant creates an abnormal splice site, causing the resulting protein to be extended. Experimental validation confirmed the prediction and identified nearby variants with similar damaging effects.
**Population-scale analysis** — Researchers analyzing whole-genome data from over 54,000 UK Biobank participants used the Atlas to uncover 22% more non-coding associations than previously detectable. By filtering to the most impactful non-coding variants, they identified 19 genomic regions associated with body mass index and pinpointed regulatory variants influencing levels of circulating proteins including PLA2G7 and EGLN1.
**Regulatory grammar mapping** — Scientists at the Stowers Institute used the motif collection to distinguish transcription factors that merely alter DNA accessibility from those that actively switch genes on or off, advancing understanding of gene regulation mechanics.
## Availability and Access
AlphaGenome Atlas is currently available as a free web portal and API for academic and non-commercial research. The underlying AlphaGenome model is also accessible on GitHub for academic use and through Model Garden on Google Cloud for commercial applications. Commercial access to the Atlas through Google Cloud is listed as forthcoming. No clinical approval or diagnostic application has been established for the resource.
## Frequently Asked Questions
**Q: What exactly is AlphaGenome Atlas?**
A: It is a precomputed catalog of molecular effect predictions for every possible single-nucleotide variant in the human genome—approximately 9 billion changes. Each entry includes regulatory predictions, a unified impact score, feature attributions, and reference to DNA motifs.
**Q: What is the AVI score?**
A: The AlphaGenome Variant Impact score is a single numerical ranking that combines regulatory predictions with protein-altering predictions. It works across both coding and non-coding regions of the genome, allowing researchers to compare variant impacts on a common scale.
**Q: How is this different from running AlphaGenome directly?**
A: Running the model on demand for each variant is computationally expensive and time-consuming. AlphaGenome Atlas eliminates this bottleneck by providing precomputed results that can be queried instantly, enabling genome-wide studies that would otherwise be impractical.
**Q: Can this be used for clinical diagnosis?**
A: The resource is currently designed for research purposes only. It has not received clinical approval, and predictions should not be used as the sole basis for medical decisions.
**Q: Who can access AlphaGenome Atlas?**
A: Academic researchers can access the Atlas through a free web portal and API. Commercial access through Google Cloud is expected to become available in the near future. The underlying AlphaGenome model is already available for both academic and commercial use through separate channels.
**Q: How large is the dataset?**
A: The Atlas occupies approximately 1 petabyte of storage, which is over 30 times larger than the AlphaFold Database of protein structure predictions.
**Q: What cell types and tissues are covered?**
A: Predictions span hundreds of human and mouse cell types and tissues, providing researchers with context-specific molecular effect data.
## Conclusion
AlphaGenome Atlas represents a significant step forward in genomics research by removing two major barriers: the computational cost of genome-scale variant analysis and the difficulty of interpreting why a variant matters. By combining precomputation with transparent feature attributions, the resource empowers researchers to move quickly from variant discovery to mechanistic understanding. Early applications in rare disease gene discovery, population genetics, and regulatory biology demonstrate the breadth of the Atlas’s utility. As access expands and the research community begins to explore this dataset at scale, further breakthroughs in our understanding of genetic variation are likely to follow.
Thank you for reading



