Skip to content

Latest commit

 

History

93 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome Chemistry Datasets Awesome

A curated list of chemistry datasets that are available for research, machine learning, and data integration. The list includes raw data sources, curated collections, and benchmarks; availability does not imply that every source permits unrestricted redistribution or commercial use.

Looking to build rather than browse? See Dataset opportunities for combinations that are practical now, common integration traps, and high-value areas that remain largely untapped.

Contributions are very welcome—please follow the guidelines and the Code of Conduct.

Contents

Text and literature

  • BC5CDR: 1,500 PubMed articles with 4,409 annotated chemicals, 5,818 diseases, and 3,116 chemical–disease interactions for named-entity recognition.
  • BioCreative V: BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions.
  • BioRxiv XML - Bulk access to the full text of bioRxiv articles for the purposes of text and data mining (TDM) is available via a dedicated Amazon S3 resource.
  • ChemTables: 788 chemical patent tables with labels of their content type. Built for semantic classification of table type. Licensed under CC BY NC 3.0.
  • Elsevier Corpus: 40,001 open-access, CC BY articles from across Elsevier journals for large-scale NLP and ML research.
  • Europe PMC - Bulk download of full text and SI of > 5 million articles.
  • IUPAC Gold Book
  • LibreText: Open-access chemistry textbook.
  • MedRxiv XML - Text and data mining is possible via dedicated Amazon S3 resource.
  • NLM Literature Archive: books, documents, and articles in life science, medicine, and healthcare; also accessible through NCBI Bookshelf. See NLMChem for 150 full-text articles with manually annotated chemical mentions.
  • OpenStax Free textbooks, including Chemistry 2e, which is released under CC-BY 4.0.
  • PubChemSTM: 281K chemical structure and text pairs
  • PubMed central: free full-text archive
  • PubMed: abstracts and outlinks
  • PubMedQA: answer research questions with yes/no/maybe using abstracts (1k expert labeled, 61.2k unlabeled and 211.3k artificially generated QA instances).
  • PubTator 3: PubMed and PMC text annotated with normalized chemicals, genes, diseases, variants, species, cell lines, and relations; available by API and bulk FTP download.
  • S2ORC: Semantic Scholar Open Research Corpus of 81.1 million English-language academic papers across many disciplines. Released under CC BY-NC 4.0.

Chemical structures and spectra

  • ChEBI: curated chemical entities, ontology terms, synonyms, structures, and cross-references; bulk SDF and database dumps are available under CC BY 4.0.
  • COCONUT: is an open source project for Natural Products (NPs) storage, search and analysis.
  • Crystallography Open Database: open-access collection of crystal structures of organic, inorganic, metal-organic compounds and minerals, excluding biopolymers. They also derived SMILES for some compounds.
  • Enamine HTS collection: 1 930 980 diverse screening compounds (37 billion molecules in 2D and 4.5 billion in 3D)
  • GDB: enumeration of molecules according to simple (feasibility and stability) rules
  • GNPS: mass spectrometry database with focus on natural products, contains untargeted (unlabelled) data.
  • MassBank: open repository of reference mass spectra with machine-readable records and structure identifiers.
  • MoNA: mass spectrometry database of real and predicted spectra for known compounds.
  • nCov-Group Data Repository: SMILES, fingerprints, descriptors, and images of millions of compounds.
  • nmrshiftdb2: is database for organic structures and their nuclear magnetic resonance (NMR) spectra.
  • PubChem: chemical structures, identifiers, properties, substances, assays, and cross-references available through APIs and bulk FTP downloads; licensing can vary with the original depositor.
  • RCSB PDB: experimentally determined 3D structures of proteins, nucleic acids, and complexes, released under CC0.
  • zinc20: ZINC20 library prepared for Deep Docking-accelerated virtual screening
  • zinc22: commercially-available compounds for virtual screening

Molecular activity and benchmarks

  • Bento: a protein-ligand docking benchmark covering rigid, flexible, de novo, blind, induced-fit, and covalent docking tasks.
  • ChEMBL: manually curated compounds, targets, assays, and bioactivity measurements, with web services and full database downloads under CC BY-SA 3.0.
  • MPCD: molecular activity prediction benchmark with 9 low-sample, narrow-scaffold inhibitor datasets and 30 higher-sample, mixed-scaffold inhibitor datasets, each visualized with TMAP.
  • MoleculeACE: a benchmark (30 HSSMS datasets in MPCD) for evaluating the predictive performance on activity cliff compounds of machine learning models.
  • PubChem BioAssay: deposited screening and assay records linked to PubChem substances and compounds; available through PUG REST and bulk downloads.

Molecular properties and benchmarks

  • ACNet: a benchmark for Activity Cliff Prediction, 400K Matched Molecular Pairs (MMPs) against 190 targets, including over 20K MMP-cliffs and 380K non-AC MMPs from ChEMBL (version 28).
  • Aquasoldb: Curation of nine open source datasets on aqueous solubility. The authors also assigned reliability groups.
  • BigSolDB 2.0: Molecular solubility in organic solvents and water in a wide range of temperatures. It contains 103944 experimentally measured solubility values of 1448 organic compounds in 213 solvents reported in the 1595 literature peer-reviewed articles. Initially BigSolDB, posted in this preprint. Recently (2025), updated to BigSolDB 2.0, published in Scientific Data.
  • BindingDB: molecular recognition database, contains 2.6M data for 1.1M Compounds and 8.10K Targets (Feb 2023)
  • BOOM: Benchmarks for Out-of-distribution Molecules is an out-of-distribution benchmark for molecular property prediction based on QM9 and LLNL-10k properties.
  • ChEBI-20: 33,010 molecule-description pairs (for molecule captioning task)
  • ESol: aqueous solubility data (log mol/L) for common organic small molecules.
  • Flashpoint: Sun et al. collected a dataset of the flashpoints of 10575 molecules from academic papers, the Gelest chemical catalogue, the DIPPR database, Lange's Handbook of Chemistry, the Hazardous Chemicals Handbook, and the PubChem database.
  • FreeSolv: Experimental and Calculated Small Molecule Hydration Free Energies
  • Harvard OPV: "experimental photovoltaic data from the literature, and corresponding quantum-chemical calculations performed over a range of geometries, each with quantum chemical results using a variety of density functionals and basis sets"
  • ILThermo: thermodynamic and transport properties of pure ionic liquids and mixtures of them.
  • Leffingwell Odor Dataset: 3523 molecules associated with expert-labeled odor descriptors from the Leffingwell PMP 2001 database
  • Lipophilicty: Experimental results of octanol/water distribution coefficient(logD at pH 7.4).
  • LLNL-10k-Dataset: DFT-calculated density and solid heat-of-formation values for approximately 10,000 CHNO molecules.
  • MD simulated monomer properties: density, cohesive energy, thermal expansion, heat of vaporization, compressibility, radius of gyration, glass transition, and diffusion constant for 410 monomers
  • MoleculeNet: benchmark suite with programmatic DeepChem loaders for molecular, reaction, image, and materials datasets.
  • oechem: On Feb 17 2023 OCHEM contained 3774118 records for 689 properties (with at least 50 records) collected from 20609 sources (user is granted a Creative Commons CC-BY (version 4.0) license to data submitted)
  • Papyrus: large-scale curated bioactivity dataset combining ChEMBL, ExCAPE-DB, and smaller public datasets.
  • minKLIFSAI: 18.8 million activity records for 452 kinases and approximately 1.2 million unique compounds (300,000 active and 900,000 inactive), collected from PubChem in January 2023.
  • Photoswitch Dataset: Curated dataset of 405 photoswitch molecules.
  • QM Datasets: QM7, QM7b, QM8, QM9, MD Trajectories
  • SolProp: Database of 1 million solvent/solute COSMO-RS calculations and 10145 experimental solvation free energies (originally published as part of this paper).
  • SOMAS: Experimental and calculated solubilities for small molecules. Originally proposed for the design of redox-flow batteries.
  • Therapeutic Data Commons: ML tasks that cover small molecules and biologics, including antibodies, peptides, miRNAs, and gene editing therapies. Original data can be found here.
  • ThermoML Archive: experimental thermophysical and thermochemical property data (in ThermoML XML format)
  • LIT-PCBA: virtual-screening benchmark with 15 target sets, 7,761 actives, and 382,674 unique inactives selected from high-confidence PubChem BioAssay data.

Target identification

  • Open Targets: is a large-scale resource that uses human genetics and genomics data for systematic drug target identification and prioritization.
  • Probes & Drugs Portal: is an interactive, open data resource for chemical biology. Overview of libraries of bioactive compounds (e.g., ChEMBL, Guide to PHARMACOLOGY), including commercial screening libraries.

Pharmacology, ADME, and metabolism

  • SIDER: drugs, adverse reactions, and indications extracted from public documents and package inserts. Released under CC BY-NC-SA 4.0.
  • Cell Effective Permeability (Caco-2) dataset: by Wang et al. is a dataset used to measure the absorption of drugs through intestinal tissue by simulating it using a human colon epithelial cancer cell line (Caco-2).
  • Clinical Trials: single zip file containing all study records (in XML) available on ClinicalTrials.gov
  • Drug–Drug–Interaction (DDI): MedLine abstracts on drug-drug interactions as well as documents describing drug-drug interactions from the DrugBank database.
  • Drug Indications Database (DID): is a dataset of structured drug-indication relations. It is intended to facilitate the building of practical, comprehensive, integrated drug ontologies.
  • EPA CompTox: is a widely used resource for chemistry, toxicity, and exposure information for hundreds of thousands of chemicals including, but not limited to, chemical properties, environmental fate, and transport, hazard, in vitro to in vivo extrapolation (IVIVE), exposure, bioactivity (each data has its license).
  • Guide to PHARMACOLOGY: is an expert-curated resource of ligand-activity-target relationships. It includes activity data even for data with unknown bioactivity value (under CC BY-SA 4.0).
  • KD-DTI: Drug-target-interaction triplets (12K training samples, 1K validation samples and 1.1K test samples). See paper.
  • KEGG PATHWAY Database(KEGG): a database resource for understanding high-level functions and utilities of the biological system, such as the cell, the organism and the ecosystem, from molecular-level information, especially large-scale molecular datasets generated by genome sequencing and other high-throughput experimental technologies.
  • LOTUS: harmonization, curation, validation and open dissemination of 750,000+ referenced structure-organism pairs (relationships between molecular structures and the living organisms from which they were identified).
  • MetXBioDB Metabolite Biotransformations: a comprehensive collection of biotransformation reactions and metabolite information from the BioTransformer database. It includes the transformation and metabolism of metabolites.
  • ONSIDES: A resource of adverse drug effects extracted from FDA structured product labels.
  • PAMPA Permeability and NCATS dataset: is a dataset of commonly employed assay to evaluate drug permeability across the cellular membrane to help in ADME prediction.
  • PsychonautWiki: catalog of mind-altering substances
  • QSAR datasets - Meta-QSAR (phase I & II): Data (extracted from ChEMBL) used in Olier et al. Meta-QSAR: a large-scale application of meta-learning to drug design and discovery.
  • State of Peptides 2026: open reference dataset of 156 peptide and peptide-adjacent compounds with legal/regulatory status buckets, categories, routes of administration, half-life, molecular weight, CAS numbers, peer-reviewed reference counts, and external knowledge-graph identifiers (PubChem/DrugBank/Wikidata). Downloadable as CSV and JSON under CC BY 4.0.
  • The Human Metabolome Database (HMDB): is a freely available electronic database containing detailed information about small molecule metabolites found in the human body.

Glycoscience

Data and registries

  • Glycan Library: almost 1,000 natural and synthetic lipid-linked, sequence-defined glycan probes, including a downloadable list of the displayed collection.
  • GlyGen: GlyGen is a data integration and dissemination project for carbohydrate and glycoconjugate related data. GlyGen retrieves information from multiple international data sources and integrates and harmonizes this data. The GlyGen web portal allows exploration of this data and execution of unique searches that cannot be performed using integrated databases in isolation. GlyGen also provides machine-readable APIs and a SPARQL endpoint to access the integrated data. Released under CC-BY-4.0 licence.
  • SugarBind: SugarBind covers knowledge of glycan binding of human pathogen lectins and adhesins. Information is collected by experts from articles published in peer-reviewed scientific journals. The data were compiled through an exhaustive search of literature published over the past 30 years by glycobiologists, microbiologists, and medical histologists.
  • UniCarb-DB: glycomics fragmentation database that stores, integrates, and processes manually annotated mass spectra.
  • GlyCosmos: is glycoscience data based on Semantic Web technology. Glycan-related data including genes, proteins, lipids, pathways, diseases and organisms. Released under CC-BY-4.0 license.
  • GlyTouCan: international glycan structure registry assigning stable accessions at resolutions from monosaccharide composition to fully defined structures. Released under CC0.
  • GlycoNAVI: is the Carbohydrate database to support carbohydrate research. Contains glycan structures, chemical synthesis, anomeric isomer proportion, NMR spectra, activity, 3D structures, and carbohydrate-protein interaction extracted from literature. Released under CC-BY licence.

Reference resources and tools

  • CAZypedia: community-driven encyclopedia of carbohydrate-active enzymes.
  • ENZYME: repository of enzyme nomenclature information.
  • ExplorEnz: interface to the approved IUBMB enzyme nomenclature and classification list.
  • GLIC: centralized software repository for glycoscientists.
  • Glycopedia: educational resource and tools for glycoscience.
  • IntEnz: integrated relational database of IUBMB enzyme nomenclature recommendations.
  • SNFG: Symbol Nomenclature for Glycans standard and drawing resources. Released under CC0.

Reactions and high-throughput screening

  • USPTO: Reactions extracted by text-mining from United States patents published between 1976 and September 2016.
  • RDB7: Computational dataset with atom-mapped SMILES, barrier heights, and reaction enthalpies calculated at CCSD(T)-F12, which is known to be very accurate. Geometries are identified via the growing string method in this paper while the high-quality energies are computed in this paper.
  • Dreher–Doyle: yields and conditions for 3,955 Pd-catalyzed Buchwald–Hartwig C–N cross-couplings.
  • Open Reaction Database: openly licensed reaction records with a detailed schema for inputs, conditions, observations, workups, outcomes, and provenance.
  • Perera: yields and conditions for 5,760 Pd-catalyzed Suzuki–Miyaura C–C cross-couplings.

Electronic laboratory notebooks

Materials and solid state

  • Digital Hydrogen Platform (DigHyd): human-validated experimental hydrogen-storage data extracted from more than 4,000 literature sources, with more than 30,000 entries. License not stated.
  • Materials Project: computed inorganic materials and molecular properties, structures, provenance, and contributed experimental data available through an API and open-data snapshots; an API key is required for the main API.
  • NOMAD: FAIR repository and archive for raw and normalized computational materials-science data with APIs and metadata-based search. Published data are available under CC BY 4.0.

Related lists

License

CC0

About

overview of datasets for ML in chemistry

Topics

Resources

Code of conduct

Contributing

Stars

424 stars

Watchers

8 watching

Forks

Releases

Packages

Used by

Contributors