Data sources

Data source inventory.

Sources
53 +20 sub
Categories
7
By category
  • Genome 15
  • Ont & Lit 13
  • Transcriptome 11
  • Proteome 10
  • Therapeutics 7
  • Visualization 3
  • Phenome 2
Key highlights
Genome
Clinical variants (ClinVar, Orphanet), structural variation (dbVar), common & rare-variant GWAS/PheWAS, fine-mapped credible sets, somatic cancer mutations, gnomAD constraint, polygenic risk scores, eQTL/cCRE overlays
Transcriptome
Bulk and single-cell expression across normal tissues, tumors, and cell lines (GTEx, TCGA, HPA, CPTAC, CELLxGENE); CRISPR/DepMap genetic dependencies; cross-species orthology (Bgee)
Proteome
Protein-protein interactions (BioGRID, InnateDB, MatrixDB, STRING), UniProt features & domains, plasma pQTLs from population cohort studies (Open Targets Platform), tumor proteomics
Phenome
UK Biobank case/control cohorts, phenotype labels, and ICD-10/EFO mappings; cross-species knockout phenotypes (IMPC)
Therapeutics
Drug-target-indication evidence (ChEMBL, IDG), pharmacogenomics (ClinPGx), drug sensitivity (PharmacoDB), antibody-drug conjugates (ADCdb), antibody sequences (PLAbDab), clinical trials
Ont & Lit
Disease ontologies (MONDO, EFO, HPO, Orphanet, ICD-10, UMLS, OncoTree), pathway/function (GO, Reactome, MSigDB, OmniPath), tissue/anatomy (UBERON), literature (EuropePMC, OpenAlex, NCBI), patents (Google Patents, EPO)
Visualization
3D protein structures (AlphaFold, PDBe, SWISS-MODEL), homology & isoform models, subcellular localization pictograms

Complete list

Data source Category Summary statistics
Cancer Hotspots Genome 2,675 variants · 224 transcripts
cBioPortal Genome 8,815,313 mutations · 495 sources
CIViC Genome clinical interpretations of variants
ClinVar Genome 553,894 variants · 29,025 phenotypes
dbVar Genome 6M+ variants
EBI eQTL Catalogue Genome 32 studies
ENCODE SCREEN cCREs Genome ~2.3M cCREs
Ensembl Genome 87,893 genes
FinnGen Genome 11,814 targets · 2,591 phenotypes
GENIE Genome 290,583 variants · 1,738 transcripts
gnomAD Genome ~17,500 transcripts
GWAS Catalog Genome 255,022 variants · 12,583 phenotypes
Open Targets Scores Genome Transcriptome Proteome Ont & Lit 11 sub-sources
Orphanet Genome Ont & Lit 2,300 phenotypes · 7,875 cross-references
PGS Catalog Genome 5,332 models · 669 traits · 781 publications
European Patent Office (EPO) Ont & Lit ~150M records
Gene Ontology (GO) Ont & Lit 40,440 terms
Google Patents Ont & Lit ~150M records
ICD-10 Ont & Lit 4,068 phenotypes
Molecular Signatures Database (MSigDB) Ont & Lit 6,269,060 gene-set memberships
NCBI Ont & Lit 95M+ gene records
OmniPath Ont & Lit 100+ resources
OncoTree Ont & Lit 834 oncotree codes · 257 cancer types
OpenAlex Ont & Lit 309M+ scholarly works
REACTOME Ont & Lit 2,863 pathway annotations
UMLS CUI Ont & Lit 726 phenotypes
International Mouse Phenotyping Consortium (IMPC) Phenome ~9,000 targets · 1,000+ phenotypes
UK Biobank Phenome 20,119 targets · 7,266 phenotypes
BioGRID Proteome 18,471 proteins · 5,840,810 interactions
InnateDB Proteome 3,737 proteins · 33,359 interactions
MatrixDB Proteome 10,423 proteins · 443,458 interactions
STRING Soon Proteome 12,174 proteins · 8,387,240 interactions
UniProt Proteome 20,779 protein-coding genes · 43,114 reviewed proteins · 1,425,900 features + domains
ADCdb Therapeutics 327 antigens · 6,500+ ADCs
ChEMBL Therapeutics 2,898,002 compounds · 1,001 targets · 2,811 indications
ClinicalTrials.gov Therapeutics 581,326 trials
ClinPGx Therapeutics 700+ compounds
Illuminating the Druggable Genome (IDG) Therapeutics 256 targets · 15,054 phenotypes
PharmacoDB Therapeutics 55,302 compounds · 5,712,751 measurements
PLAbDab Therapeutics 150,000+ antibody sequences
Bgee Transcriptome 60,490 genes · 3 species
CELLxGENE Discover Transcriptome 33M+ cells · 436 datasets · 2,700+ cell types
CPTAC (pancan) Transcriptome Proteome 104,235 transcripts · 1,565 samples · 10 datasets
Gene Expression Omnibus (GEO) Soon Transcriptome Proteome 273,310 series · 4,352 datasets
GTEx Transcriptome 58,988 transcripts · 68 tissues
Human Protein Atlas Transcriptome Proteome 20,141 transcripts · 59 tissues
PanglaoDB Transcriptome 178 cell types
Prostate Cancer Atlas Transcriptome Proteome 92,144 samples
TCGA Transcriptome 17,102 transcripts · 13,454 samples · 34 datasets
TCGA/TARGET/GTEx (UCSC XenaBrowser) Transcriptome 17,409 transcripts · 28,823 samples · 37 datasets
SWISS-MODEL Visualization 3,939,534 models · 234,439 structures · 12 proteomes
SwissBioPic Soon Visualization ~50 species with subcellular pictograms
3D Beacons Visualization 9 sub-sources

Data attributions

Kasvu Discovery integrates data from a wide range of publicly available scientific databases and resources. We are grateful to the teams and communities behind each of these efforts. Where required by licence terms, we provide the following attributions.

UK Biobank
Genomic and phenotypic summary statistics are sourced via the Pan-UK Biobank project and used in accordance with the UK Biobank terms of access. Karczewski, Gupta, Kanai et al., "Pan-UK Biobank GWAS improves discovery, analysis of genetic architecture, and resolution into ancestry-enriched effects." medRxiv. 2024. doi: 10.1101/2024.03.13.24303864.
gnomAD
Data from the Genome Aggregation Database (gnomAD). Karczewski et al., "The mutational constraint spectrum quantified from variation in 141,456 humans." Nature. 2020. doi: 10.1038/s41586-020-2308-7.
Gene Ontology
Gene Ontology data and data products are licensed under the Creative Commons Attribution 4.0 International licence. Ashburner et al., "Gene Ontology: tool for the unification of biology." Nature Genetics. 2000. doi: 10.1038/75556. The Gene Ontology Consortium, "The Gene Ontology knowledgebase in 2023." Genetics. 2023. doi: 10.1093/genetics/iyad031.
GTEx
Data were obtained from the GTEx Portal. The Genotype-Tissue Expression (GTEx) Project was supported by the Common Fund of the Office of the Director of the National Institutes of Health.
TCGA
Data were generated by The Cancer Genome Atlas Research Network. We acknowledge the TCGA Research Network and the study participants who contributed to this resource.
Human Protein Atlas
Data from the Human Protein Atlas, available at proteinatlas.org. Uhlén et al., "Proteomics: Tissue-based map of the human proteome." Science. 2015. doi: 10.1126/science.1260419.
ChEMBL
Data sourced from ChEMBL (EMBL-EBI), made available under the Creative Commons Attribution-ShareAlike 3.0 Unported licence (CC BY-SA 3.0).
FinnGen
We acknowledge the FinnGen study. Kurki et al., "FinnGen provides genetic insights from a well-phenotyped isolated population." Nature. 2023. doi: 10.1038/s41586-022-05473-8.
ICD-10
ICD-10 codes are copyright of the World Health Organization. Source: World Health Organization. Used under the Creative Commons Attribution-NoDerivs 3.0 IGO licence (CC BY-ND 3.0 IGO).
Reactome
Pathway data sourced from Reactome, a free, open-source, curated and peer-reviewed pathway database. Gillespie et al., "The reactome pathway knowledgebase 2022." Nucleic Acids Research. 2022. doi: 10.1093/nar/gkab1028.
UMLS
UMLS terminology data sourced from the Unified Medical Language System (UMLS), National Library of Medicine, National Institutes of Health.

All third-party database names and trademarks are the property of their respective owners. Listing a data source on this page does not imply endorsement by, or affiliation with, the providing organisation.