• Home
  • Search
  • SNP interaction detection with Random Forests in high-dimensional genetic data
  • Cite Icon105
  • https://doi.org/10.1186/1471-2105-13-164Copy DOI Icon

SNP interaction detection with Random Forests in high-dimensional genetic data

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

BackgroundIdentifying variants associated with complex human traits in high-dimensional data is a central goal of genome-wide association studies. However, complicated etiologies such as gene-gene interactions are ignored by the univariate analysis usually applied in these studies. Random Forests (RF) are a popular data-mining technique that can accommodate a large number of predictor variables and allow for complex models with interactions. RF analysis produces measures of variable importance that can be used to rank the predictor variables. Thus, single nucleotide polymorphism (SNP) analysis using RFs is gaining popularity as a potential filter approach that considers interactions in high-dimensional data. However, the impact of data dimensionality on the power of RF to identify interactions has not been thoroughly explored. We investigate the ability of rankings from variable importance measures to detect gene-gene interaction effects and their potential effectiveness as filters compared to p-values from univariate logistic regression, particularly as the data becomes increasingly high-dimensional.ResultsRF effectively identifies interactions in low dimensional data. As the total number of predictor variables increases, probability of detection declines more rapidly for interacting SNPs than for non-interacting SNPs, indicating that in high-dimensional data the RF variable importance measures are capturing marginal effects rather than capturing the effects of interactions.ConclusionsWhile RF remains a promising data-mining technique that extends univariate methods to condition on multiple variables simultaneously, RF variable importance measures fail to detect interaction effects in high-dimensional data in the absence of a strong marginal component, and therefore may not be useful as a filter technique that allows for interaction effects in genome-wide data.

Loading PDF

Similar Papers
  • PDF
  • Research Article
  • Citations3603

Bias in random forest variable importance measures: illustrations, sources and a solution.

  • Jan 25, 2007
  • BMC Bioinformatics
  • Carolin Strobl +3
  • Research Article
  • Citations20

An integrated approach to reduce the impact of minor allele frequency and linkage disequilibrium on variable importance measures for genome-wide data

  • Jul 30, 2012
  • Bioinformatics
  • Raymond Walters +2
  • Peer Review Report

Editor's evaluation: Derivation and external validation of clinical prediction rules identifying children at risk of linear growth faltering

  • Sep 05, 2022
  • Eduardo Franco
  • Conference Article
  • Citations10

Detection of SNP-SNP Interactions in Genome-wide Association Data Using Random Forests and Association Rules

  • Dec 01, 2018
  • Tung Nguyen +1
  • Conference Article

Optimization of random forest algorithm based on mixed sampling additional feature selection

  • Jan 06, 2023
  • Haobo Cui +2
  • Research Article
  • Citations102

A weighted random forests approach to improve predictive performance

  • Jul 08, 2013
  • Statistical Analysis and Data Mining: The ASA Data Science Journal
  • Stacey J Winham +2
  • Research Article
  • Citations190

Improving land cover classification in an urbanized coastal area by random forests: The role of variable selection

  • Sep 21, 2020
  • Remote Sensing of Environment
  • Fang Zhang +1
  • Research Article
  • Citations22

Pathway-based identification of SNPs predictive of survival

  • Feb 02, 2011
  • European Journal of Human Genetics
  • Herbert Pang +2
  • Research Article
  • Citations2

The Impact of Variable Omission on Variable Importance Measures of Cart, Random Forest, and Boosting Algorithms

  • Mar 30, 2022
  • Journal of Statistical Research
  • W Holmes Finch
  • PDF
  • Research Article
  • Citations6

Assessing the Influence of Operational Variables on Process Performance in Metallurgical Plants by Use of Shapley Value Regression

  • Oct 22, 2022
  • Metals
  • Xiu Liu +1
  • Book Chapter

Supervised Machine Learning: Application Example Using Random Forest in R

  • Aug 08, 2019
  • Bharatendra Rai
  • PDF
  • Research Article
  • Citations673

Evaluation of variable selection methods for random forests and omics data sets.

  • Oct 16, 2017
  • Briefings in bioinformatics
  • Frauke Degenhardt +2
  • Research Article
  • Citations111

Nonparametric variable importance assessment using machine learning techniques.

  • Dec 08, 2020
  • Biometrics
  • Brian D Williamson +3
  • Research Article
  • Citations86

Identification of important factors in an inpatient fall risk prediction model to improve the quality of care using EHR and electronic administrative data: A machine-learning approach

  • Sep 15, 2020
  • International journal of medical informatics
  • David S Lindberg +11
  • Book Chapter
  • Citations1

Random Forests for Survival Analysis and High-Dimensional Data

  • Jan 01, 2023
  • Ruoqing Zhu +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.