• Home
  • Search
  • Towards Robust Performance Guarantees for Models Learned from High-Dimensional Data
  • Open Access IconOpen Access
  • Cite Icon2
  • https://doi.org/10.1007/978-3-319-11056-1_3Copy DOI Icon

Towards Robust Performance Guarantees for Models Learned from High-Dimensional Data

  • Jan 1, 2015
  • Rui Henriques +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Models learned from high-dimensional spaces, where the high number of features can exceed the number of observations, are susceptible to overfit since the selection of subspaces of interest for the learning task is prone to occur by chance. In these spaces, the performance of models is commonly highly variable and dependent on the target error estimators, data regularities and model properties. High-variable performance is a common problem in the analysis of omics data, healthcare data, collaborative filtering data, and datasets composed by features extracted from unstructured data or mapped from multi-dimensional databases. In these contexts, assessing the statistical significance of the performance guarantees of models learned from these high-dimensional spaces is critical to validate and weight the increasingly available scientific statements derived from the behavior of these models. Therefore, this chapter surveys the challenges and opportunities of evaluating models learned from big data settings from the less-studied angle of big dimensionality. In particular, we propose a methodology to bound and compare the performance of multiple models. First, a set of prominent challenges is synthesized. Second, a set of principles is proposed to answer the identified challenges. These principles provide a roadmap with decisions to: i) select adequate statistical tests, loss functions and sampling schema, ii) infer performance guarantees from multiple settings, including varying data regularities and learning parameterizations, and iii) guarantee its applicability for different types of models, including classification and descriptive models. To our knowledge, this work is the first attempt to provide a robust and flexible assessment of distinct types of models sensitive to both the dimensionality and size of data. Empirical evidence supports the relevance of these principles as they offer a coherent setting to bound and compare the performance of models learned in high-dimensional spaces, and to study and refine the behavior of these models.Keywordshigh-dimensional dataperformance guaranteesstatistical significance of learning modelserror estimatorsclassificationbiclustering

Similar Papers
  • PDF
  • Research Article
  • Citations56

Pushing the limits of solubility prediction via quality-oriented data selection.

  • Dec 17, 2020
  • iScience
  • Murat Cihan Sorkun +2
  • Book Chapter
  • Citations17

Multivariate Statistical Methods for High-Dimensional Multiset Omics Data Analysis

  • Nov 01, 2019
  • Attila Csala +1
  • Conference Article
  • Citations15

Scalability and Total Recall with Fast CoveringLSH

  • Oct 24, 2016
  • Ninh Pham +1
  • Book Chapter
  • Citations9

Pairwise Constrained Clustering for Sparse and High Dimensional Feature Spaces

  • Jan 01, 2009
  • Su Yan +3
  • Research Article
  • Citations9

Graphs as navigational infrastructure for high dimensional data spaces

  • Feb 11, 2011
  • Computational Statistics
  • C B Hurley +1
  • PDF
  • Research Article
  • Citations15

Finding the best trade-off between performance and interpretability in predicting hospital length of stay using structured and unstructured data.

  • Nov 30, 2023
  • PloS one
  • Franck Jaotombo +3
  • Dissertation

Enhanced Decision Maps for Exploring Classification Models

  • Jun 16, 2025
  • Yu Wang
  • Book Chapter

Efficient Streaming Detection of Hidden Clusters in Big Data Using Subspace Stream Clustering

  • Jan 01, 2014
  • Marwan Hassani +1
  • PDF
  • Research Article
  • Citations38

Implementing Magnetic Resonance Imaging Brain Disorder Classification via AlexNet–Quantum Learning

  • Jan 10, 2023
  • Mathematics
  • Naif Alsharabi +3
  • Research Article
  • Citations23

Investigating the distribution of archaeological sites: Multiparametric vs probability models and potentials for remote sensing data

  • Apr 25, 2018
  • Applied Geography
  • Mariangela Noviello +4
  • Conference Article
  • Citations1

On the Necessary and Sufficient Conditions of a Meaningful Distance Function for High Dimensional Data Space

  • Apr 20, 2006
  • Chih-Ming Hsu +1
  • Research Article
  • Citations3

Applying machine learning to high-dimensional proteomics datasets for the identification of Alzheimer’s disease biomarkers

  • Mar 03, 2025
  • Fluids and Barriers of the CNS
  • Christoffer Ivarsson Orrelid +7
  • Research Article
  • Citations9

Multiplicative distance: a method to alleviate distance instability for high-dimensional data

  • Dec 28, 2014
  • Knowledge and Information Systems
  • Jafar Mansouri +1
  • Book Chapter

Using Distribution Divergence to Predict Changes in the Performance of Clinical Predictive Models

  • Jan 01, 2021
  • Mohammadamin Tajgardoon +1
  • Conference Article
  • Citations106

Color and Position versus Texture Features for Endoscopic Polyp Detection

  • May 01, 2008
  • Lu Alexandre +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.