• Home
  • Search
  • Data, Machine Learning, and Human Domain Experts: None Is Better than Their Collaboration
  • Cite Icon29
  • https://doi.org/10.1080/10447318.2021.2002040Copy DOI Icon

Data, Machine Learning, and Human Domain Experts: None Is Better than Their Collaboration

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

ABSTRACT Software designers are coming up with tools having machine learning (ML) capabilities integrated into them. Even naive users can build ML-based models for their respective problem domains using these tools. However, a majority of the advanced ML models behave like a black box in the sense that their behavior is not easily interpretable by humans. This lack of human interpretability acts as a hurdle toward acceptance and subsequent deployment of ML-based solutions. This research work is based on an intuitive idea that in an ideal ML-based solution, the provided dataset, the ML model learned using this dataset, and human domain experts must agree in terms of their perception regarding the problem domain. The objective of this work is to propose a framework that first listens to the dataset, the ML model, and human domain experts individually, and then measure the degree of agreement between them. Listening to the dataset refers identifying important characteristics by computing information gain using entropy and Gini index. Listening to the ML model refers interpreting its decision-making behavior by computing variable importance measures. Listening to human experts refers taking feedback in terms of important features as per their perception regarding the problem domain. The task of measuring agreement between dataset, model, and human experts has been modeled as a “two-judges n-participants” rank correlation problem. The proposed approach has been evaluated on a novel problem domain of predicting joining behavior of freshmen students. A positive degree of agreement was observed between the dataset, the ML model, and human experts in terms of their perception regarding the problem domain. The agreement between the provided dataset, the learned model, and human domain experts helps verify learning acquired by the ML model against the prevailing domain knowledge. This work has the potential to form the basis for developing formal quantitative metrics for evaluating ML models in terms of reliable learning and capability to facilitate trust of human users.

Similar Papers
  • Research Article
  • Citations19

Data reformation – A novel data processing technique enhancing machine learning applicability for predicting streamflow extremes

  • Nov 03, 2023
  • Advances in Water Resources
  • Vinh Ngoc Tran +2
  • Research Article
  • Citations42

Transfer learning enables prediction of steel corrosion in concrete under natural environments

  • Feb 24, 2024
  • Cement and Concrete Composites
  • Haodong Ji +3
  • Research Article
  • Citations15

Enhancing Coronary Artery Disease Prognosis: A Novel Dual-Class Boosted Decision Trees Strategy for Robust Optimization

  • Jan 01, 2024
  • IEEE Access
  • Tariq Mahmood +6
  • Research Article
  • Citations1

Machine Learning for Intensive Care Unit Length-of-Stay Prediction: A Simulation-Based Approach to Bed Capacity Management.

  • Dec 26, 2025
  • Medical decision making : an international journal of the Society for Medical Decision Making
  • Sara Garber +1
  • Research Article
  • Citations2

A systematic literature review on Machine Learning Model evaluation on healthcare applications

  • Jun 14, 2023
  • Research, Society and Development
  • Cezar Miranda Paula De Souza +9
  • PDF
  • Research Article
  • Citations4

A genome-wide association study coupled with machine learning approaches to identify influential demographic and genomic factors underlying Parkinson’s disease

  • Sep 29, 2023
  • Frontiers in Genetics
  • Md Asad Rahman +1
  • Conference Article
  • Citations4

Predicting upper limb disability progression in primary progressive multiple sclerosis using machine learning and statistical methods

  • Dec 09, 2021
  • Sally Mostafa +7
  • PDF
  • Research Article
  • Citations40

Machine Learning Models for Blood Glucose Level Prediction in Patients With Diabetes Mellitus: Systematic Review and Network Meta-Analysis.

  • Nov 20, 2023
  • JMIR Medical Informatics
  • Kui Liu +9
  • PDF
  • Research Article
  • Citations6

Simple Linear Cancer Risk Prediction Models With Novel Features Outperform Complex Approaches.

  • May 01, 2022
  • JCO Clinical Cancer Informatics
  • Scott Kulm +3
  • PDF
  • Research Article
  • Citations42

Prediction of shear behavior of glass FRP bars-reinforced ultra-highperformance concrete I-shaped beams using machine learning

  • Aug 30, 2023
  • International Journal of Mechanics and Materials in Design
  • Asif Ahmed +6
  • Research Article

Machine learning for predicting outcomes, complications and resource utilisation after hip arthroscopy: A systematic review.

  • Jan 31, 2026
  • Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA
  • Madison Lee +4
  • Research Article

Enhanced Language Models for Predicting and Understanding HIV Care Disengagement: A Case Study in Tanzania

  • May 08, 2025
  • Research Square
  • Waverly Wei +16
  • Research Article
  • Citations527

Systematic literature review of machine learning based software development effort estimation models

  • Sep 16, 2011
  • Information and Software Technology
  • Jianfeng Wen +4
  • PDF
  • Research Article
  • Citations27

Development of Monthly Reference Evapotranspiration Machine Learning Models and Mapping of Pakistan—A Comparative Study

  • May 23, 2022
  • Water
  • Jizhang Wang +8
  • Research Article
  • Citations1

Do You Consent to the Use of Your Biological Data for Training ML and AI Models? Online Survey Targeting Clinicians and Researchers.

  • Jan 27, 2024
  • Web3 Journal: ML in Health Science
  • Yury Rusinovich +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.