Over the past two decades, quantitative structure–activity relationship (QSAR) modeling has evolved substantially, driven by improved data accessibility, open-source descriptor generation, mature machine learning libraries, and scalable cloud computing. Large-scale benchmarking studies using public datasets have demonstrated the feasibility of building predictive models across hundreds of endpoints. In parallel, automated machine learning (Auto-ML) approaches have emerged as a promising means to lower the barrier to QSAR model development, enabling competitive performance without extensive expert intervention. Here, we describe the design and implementation of an automated QSAR modeling system integrated into the CDD Vault platform, referred to as CDD Vault Inference Models. The system automatically trains, evaluates, and deploys regression models whenever new assay data become available, without requiring users to select endpoints, descriptors, or learning algorithms. Using public datasets from ChEMBL, we developed a fully automated workflow for model training and continuous evaluation. Models are released when a conservative performance threshold is achieved. The system is currently focused on building regression models. To give users a handle on model uncertainty, we also provide conformal prediction intervals. We discuss the implications of deploying fully automated QSAR models in a production environment and outline future extensions. Together, this work demonstrates that automated, continuously updated QSAR modeling can provide practical and scalable decision support for drug discovery, particularly in settings where dedicated modeling expertise is limited.
Natural products (NPs) continue to inform the discovery and development of a diversity of drugs, both marketed and investigational. Pain, one of the most common of human experiences and profound challenges in medicine and biology, has emerged at the core of an urgent societal problem, in the United States and globally. The present study employs a retrospective analysis of an extensive set of published literature curated in the NAPRALERT database to identify NPs with experimental evidence of bioactivity supporting the selection and prioritization of NP leads with promise in pain management. The NAPRALERT pain data set currently documents >38,000 pain-relevant experiments reported in >1,750 distinct journals. The evidence presented here was annotated from >10,000 distinct scientific publications identifying NP extracts and isolates with experimental biological data indicating positive mitigation of pain, inflammation, and/or modulation of nociceptive signaling targets. Correlation of ethnomedical uses with experimental data represents a value-added approach to the selection and prioritization of leads. Dissemination of this unique NP/pain data set, with experimental data and information applicable to basic, translational, and clinical science stakeholders alike, furnishes practical evidence in support of a rational selection of NPs for directed pain research. A large portion of the NAPRALERT pain-relevant data set, along with a set of query tools designed to assist user-directed selection and prioritization of leads, are presented as Supporting Information in order to mitigate the limitations inherent in presenting such a large data set in (print) format. To support user efforts, this report involves explication of NAPRALERT data organization and the articulation of rational approaches to user-guided selection of evidence-based NP leads.