• Home
  • Search
  • Comparing massively-multitask regression algorithms for drug discovery.
  • Cite Icon1
  • https://doi.org/10.1007/s10822-026-00761-1Copy DOI Icon

Comparing massively-multitask regression algorithms for drug discovery.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Massively-multitask regression models (MMRMs) have revolutionized activity prediction for drug discovery. MMRMs trained on millions of compounds and many thousands of assays can predict bioactivity with accuracy comparable to 4-concentration IC50 experiments. This report compares six MMRMs: pQSAR, Alchemite, MT-DNN, MetaNN, Macau and IMC. Models were trained by experts in each method, on identical sets of 159 kinase and 4276 diverse ChEMBL assays, employing realistically novel training/test set splits. Results were compared both qualitatively and with statistical rigor. Our use-case is imputing full bioactivity profiles for the very sparse compound collections on which the models were trained. MMRMs performed much better than the single-task random forest regression (ST-RFR) model. Five MMRMs train all models simultaneously, so must leave out test-set measurements from all assays to avoid leakage (here 25% of data), whereas one method trains models one-at-a-time, so only holds out test data for that assay (< 1% of data). Thus, all algorithms were compared both using 75/25 splits, and when possible, 99 + / < 1 splits. Many MMRM evaluations achieved similar accuracy when tested on the same split. However, when evaluated on 75/25 splits, all MMRMs performed much worse than when evaluated on 99 + / < 1% splits. Thus, while many MMRMs produce comparable final production models (trained on all the data), models that require 75/25 splits greatly underestimate the accuracy of the final models. While outstanding for imputations, MMRMs proved little better than ST-RFR for compounds very unlike the training collection. Thus, MMRMs are best for hit-finding, off-target, promiscuity, MoA, polypharmacology or drug-repurposing within the training collection. Since accuracy is not a deciding factor, other pros and cons of each method are also described.

Similar Papers
  • Peer Review Report

Editor's evaluation: Derivation and external validation of clinical prediction rules identifying children at risk of linear growth faltering

  • Sep 05, 2022
  • Eduardo Franco
  • Research Article

Examining the utility of digital phenotyping for the prediction of intrusive experiences

  • Dec 31, 2026
  • European Journal of Psychotraumatology
  • Tomas Meaney +3
  • Conference Article
  • Citations3

A Stable Passenger Flow Forecast Approach for Newly Opened Metro Stations Based on Multi-Source Data and Random Forest Regression Model

  • Oct 21, 2022
  • Kang Yao +4
  • Research Article
  • Citations180

Estimation of heavy metals using deep neural network with visible and infrared spectroscopy of soil

  • Jun 18, 2020
  • Science of The Total Environment
  • Jongcheol Pyo +4
  • Research Article
  • Citations34

Interpreting Highly Variable Indoor PM2.5 in Rural North China Using Machine Learning.

  • May 08, 2023
  • Environmental science & technology
  • Yatai Men +9
  • Research Article

Mechanistic insights into sulfamethoxazole removal in activated sludge systems through machine learning and microcosm experiments.

  • Mar 01, 2026
  • Water research
  • Junjie Chen +6
  • PDF
  • Research Article
  • Citations19

Evaluating how lodging affects maize yield estimation based on UAV observations.

  • Jan 17, 2023
  • Frontiers in Plant Science
  • Yuan Liu +11
  • Research Article
  • Citations16

Comparison of random regression and repeatability models to predict breeding values from test-day records of Norwegian goats

  • Jan 26, 2013
  • Journal of Dairy Science
  • S Andonov +5
  • Research Article
  • Citations54

Estimating monthly average temperature by remote sensing in China

  • Jan 09, 2019
  • Advances in Space Research
  • Long Li +1
  • PDF
  • Research Article
  • Citations27

Using machine learning to derive cloud condensation nuclei number concentrations from commonly available measurements

  • Nov 05, 2020
  • Atmospheric Chemistry and Physics
  • Arshad Arjunan Nair +1
  • PDF
  • Research Article
  • Citations33

Analysis of Financing Efficiency of Chinese Agricultural Listed Companies Based on Machine Learning

  • Jan 01, 2019
  • Complexity
  • Lixia Liu +1
  • Research Article
  • Citations8

Synergy of optical and synthetic aperture radar data for early-stage crop yield estimation: a case study over a state of Germany

  • Feb 08, 2022
  • Geocarto International
  • Parmita Ghosh +5
  • Research Article
  • Citations3

Characterization of naturally occurring radioactive material dynamics in community water systems using groundwater from Ganghwa Island, Republic of Korea

  • Nov 20, 2023
  • Journal of Hydrology
  • Eunhyung Lee +9
  • Research Article

Use of statistical models for predicting oral health status of children with cerebral palsy in Sri Lanka

  • Feb 28, 2021
  • Biometrics &amp; Biostatistics International Journal
  • H.B.W.M.D.M Weerasekara +2
  • Research Article
  • Citations50

Towards safer streets: A framework for unveiling pedestrians’ perceived road safety using street view imagery

  • Nov 28, 2023
  • Accident Analysis &amp; Prevention
  • Omar Faruqe Hamim +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.