• Home
  • Search
  • Evaluation of variable selection methods for random forests and omics data sets.
  • Cite Icon673
  • https://doi.org/10.1093/bib/bbx124Copy DOI Icon

Evaluation of variable selection methods for random forests and omics data sets.

Show More
  • Abstract
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Machine learning methods and in particular random forests are promising approaches for prediction based on high dimensional omics data sets. They provide variable importance measures to rank predictors according to their predictive power. If building a prediction model is the main goal of a study, often a minimal set of variables with good prediction performance is selected. However, if the objective is the identification of involved variables to find active networks and pathways, approaches that aim to select all relevant variables should be preferred. We evaluated several variable selection procedures based on simulated data as well as publicly available experimental methylation and gene expression data. Our comparison included the Boruta algorithm, the Vita method, recurrent relative variable importance, a permutation approach and its parametric variant (Altmann) as well as recursive feature elimination (RFE). In our simulation studies, Boruta was the most powerful approach, followed closely by the Vita method. Both approaches demonstrated similar stability in variable selection, while Vita was the most robust approach under a pure null model without any predictor variables related to the outcome. In the analysis of the different experimental data sets, Vita demonstrated slightly better stability in variable selection and was less computationally intensive than Boruta.In conclusion, we recommend the Boruta and Vita approaches for the analysis of high-dimensional data sets. Vita is considerably faster than Boruta and thus more suitable for large data sets, but only Boruta can also be applied in low-dimensional settings.

Loading PDF

Similar Papers
  • Research Article
  • Citations52

Distributed feature selection (DFS) strategy for microarray gene expression data to improve the classification performance

  • Apr 27, 2018
  • Clinical Epidemiology and Global Health
  • Sai Prasad Potharaju +1
  • Research Article
  • Citations23

Evaluation of the effect of chance correlations on variable selection using Partial Least Squares-Discriminant Analysis

  • Aug 09, 2013
  • Talanta
  • Julia Kuligowski +6
  • Research Article
  • Citations9

High-dimensional sparse vine copula regression with application to genomic prediction.

  • Jan 29, 2024
  • Biometrics
  • Özge Sahin +1
  • Research Article
  • Citations75

DBFS: An effective Density Based Feature Selection scheme for small sample size and high dimensional imbalanced data sets

  • Aug 17, 2012
  • Data & Knowledge Engineering
  • Mina Alibeigi +2
  • PDF
  • Research Article
  • Citations6

A Novel Density-based Technique for Outlier Detection of High Dimensional Data Utilizing Full Feature Space

  • Mar 25, 2021
  • Information Technology and Control
  • Mujeeb Ur Rehman +1
  • Research Article
  • Citations26

Eigenvectors from Eigenvalues Sparse Principal Component Analysis (EESPCA)

  • Oct 01, 2021
  • Journal of Computational and Graphical Statistics
  • H Robert Frost
  • Research Article
  • Citations2

Differential privacy protection algorithm for large data sources based on normalized information entropy Bayesian network

  • Aug 01, 2024
  • Journal of Physics: Conference Series
  • Guangyuan Ni +1
  • PDF
  • Research Article
  • Citations130

Recursive cluster elimination (RCE) for classification and feature selection from gene expression data.

  • May 02, 2007
  • BMC Bioinformatics
  • Malik Yousef +3
  • PDF
  • Research Article
  • Citations31

Genome-wide prediction of transcriptional regulatory elements of human promoters using gene expression and promoter analysis data

  • Jul 04, 2006
  • BMC Bioinformatics
  • Seon-Young Kim +1
  • Conference Article
  • Citations5

A Comparative Study for Classification on Different Domain

  • Feb 26, 2018
  • Noviyanti Tri Maretta Sagala +1
  • Book Chapter
  • Citations67

Bayesian Models for Sparse Regression Analysis of High Dimensional Data*

  • Oct 06, 2011
  • Sylvia Richardson +2
  • Research Article
  • Citations19

Improved Sparse Multi-Class SVM and Its Application for Gene Selection in Cancer Classification

  • Jan 01, 2013
  • Cancer Informatics
  • Lingkang Huang +3
  • Research Article
  • Citations80

Speeding up k-Means algorithm by GPUs

  • May 08, 2012
  • Journal of Computer and System Sciences
  • You Li +3
  • PDF
  • Research Article
  • Citations241

Class prediction for high-dimensional class-imbalanced data

  • Oct 20, 2010
  • BMC Bioinformatics
  • Rok Blagus +1
  • Book Chapter
  • Citations2

On the Effectiveness of Dimensionality Reduction Techniques on High Dimensionality Datasets

  • Jan 01, 2023
  • Salah Eddine Henouda +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.