• Home
  • Search
  • Feature selection to increase the random forest method performance on high dimensional data
  • Cite Icon28
  • https://doi.org/10.26555/ijain.v6i3.471Copy DOI Icon

Feature selection to increase the random forest method performance on high dimensional data

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Random Forest is a supervised classification method based on bagging (Bootstrap aggregating) Breiman and random selection of features. The choice of features randomly assigned to the Random Forest makes it possible that the selected feature is not necessarily informative. So it is necessary to select features in the Random Forest. The purpose of choosing this feature is to select an optimal subset of features that contain valuable information in the hope of accelerating the performance of the Random Forest method. Mainly for the execution of high-dimensional datasets such as the Parkinson, CNAE-9, and Urban Land Cover dataset. The feature selection is done using the Correlation-Based Feature Selection method, using the BestFirst method. Tests were carried out 30 times using the K-Cross Fold Validation value of 10 and dividing the dataset into 70% training and 30% testing. The experiments using the Parkinson dataset obtained a time difference of 0.27 and 0.28 seconds faster than using the Random Forest method without feature selection. Likewise, the trials in the Urban Land Cover dataset had 0.04 and 0.03 seconds, while for the CNAE-9 dataset, the difference time was 2.23 and 2.81 faster than using the Random Forest method without feature selection. These experiments showed that the Random Forest processes are faster when using the first feature selection. Likewise, the accuracy value increased in the two previous experiments, while only the CNAE-9 dataset experiment gets a lower accuracy. This research’s benefits is by first performing feature selection steps using the Correlation-Base Feature Selection method can increase the speed of performance and accuracy of the Random Forest method on high-dimensional data.

Similar Papers
  • PDF
  • Research Article
  • Citations135

Random KNN feature selection - a fast and stable alternative to Random Forests

  • Nov 18, 2011
  • BMC Bioinformatics
  • Shengqiao Li +2
  • Conference Article
  • Citations5

Classification Application Based on Mutual Information and Random Forest Method for High Dimensional Data

  • Aug 01, 2017
  • Qingqing Kong +3
  • Research Article
  • Citations19

Image classification method of cashmere and wool based on the multi-feature selection and random forest method

  • Sep 22, 2021
  • Textile Research Journal
  • Yaolin Zhu +3
  • Conference Article

Optimization of random forest algorithm based on mixed sampling additional feature selection

  • Jan 06, 2023
  • Haobo Cui +2
  • Research Article
  • Citations2

Effective hybrid feature subset selection for multilevel datasets using decision tree classifiers

  • Jan 01, 2023
  • International Journal of Advanced Intelligence Paradigms
  • S Dinakaran +1
  • Book Chapter

Machine Learning and Optimization Techniques for Cyberattack Prevention in Wireless Networks

  • Feb 28, 2025
  • Usharani Bhimavarapu
  • PDF
  • Research Article
  • Citations106

Unbiased Feature Selection in Learning Random Forests for High-Dimensional Data

  • Jan 01, 2015
  • The Scientific World Journal
  • Thanh-Tung Nguyen +2
  • Peer Review Report

Decision letter: Applying causal discovery to single-cell analyses using CausalCell

  • Aug 14, 2022
  • Babak Momeni
  • Preprint Article
  • Citations2

An assessment of Random Forest wrappers for selecting important features of spectroscopy data in the modelling of soil properties

  • Mar 27, 2022
  • Francisco M Canero +3
  • PDF
  • Research Article
  • Citations18

Feature Selection and Feature Stability Measurement Method for High-Dimensional Small Sample Data Based on Big Data Technology.

  • Jan 01, 2021
  • Computational Intelligence and Neuroscience
  • Chengyuan Huang
  • PDF
  • Research Article
  • Citations5

Predicting Coronary Stenosis Progression Using Plaque Fatigue From IVUS-Based Thin-Slice Models: A Machine Learning Random Forest Approach

  • May 10, 2022
  • Frontiers in Physiology
  • Xiaoya Guo +8
  • Research Article
  • Citations9

Feature genes predicting the FLT3/ITD mutation in acute myeloid leukemia

  • May 12, 2016
  • Molecular Medicine Reports
  • Chenglong Li +3
  • PDF
  • Research Article
  • Citations19

Predicting A-to-I RNA editing by feature selection and random forest.

  • Oct 22, 2014
  • PLoS ONE
  • Yang Shu +4
  • Research Article
  • Citations1

Design of feature selection algorithm for high-dimensional network data based on supervised discriminant projection.

  • Jun 26, 2023
  • PeerJ. Computer science
  • Zongfu Zhang +4
  • Research Article
  • Citations43

Determining the Extinguishing Status of Fuel Flames With Sound Wave by Machine Learning Methods

  • Jan 01, 2021
  • IEEE Access
  • Murat Koklu +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.