The Genomics Data Lake contains links, documentations and access to multiple public genomics datasets that can be accessed for free and integrated into genomics analysis workflows and applications. The datasets can include genome sequences, variant info, and subject/sample metadata in BAM, FASTA, VCF, CSV, and/or PARQUET file formats. The Data Lake contains data from: Illumina, Genome Reference Consortium, ClinVar, SpnEff, Genome Aggregation Database (gnomAD), 1000 Genomes Project, OpenCRAVAT, Encyclopedia of DNAElements (ENCODE) Consortium, GATK resource bundle, The Cancer Genome Atlas (TCGA), Pan-ancestry genetic analysis of the UK Biobank (Pan-UKBB). As of May 2025, links to the different datasets is being deprecated and the home page has been removed. However, information about the data and sources is still accessible in the dataset catalog, which one page for each source.
Microsoft no longer provides access to the catalog. The image below represents an archived version of when the catalog was accessible.
| Identifier | URL: https://learn.microsoft.com/en-us/azure/open-datasets/dataset-genomics-data-lake |
| Type | Catalog |
| Topics/About | |
| Language | English |
| Accessible For Free | TRUE |
| Biological Scale | population, molecular |
| Geographical Scope | Earth |
| Collection End Date | 2025/05/31 |
| File Format | CSV, PARQUET, FASTA |
| Version | unspecified |
| HTTP Status Code | NA |
| HTTP Checked On | NA |