• Home
  • Search
  • Gene selection stability's dependence on dataset difficulty
  • Cite Icon12
  • https://doi.org/10.1109/iri.2013.6642491Copy DOI Icon

Gene selection stability's dependence on dataset difficulty

  • Aug 1, 2013
  • David J Dittman +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Identifying important biomarkers to improve disease diagnosis and treatment is a significant topic of research in bioinformatics. However, bioinformatics datasets frequently have a large number of features per sample or instance. This problem, known as “high dimensionality,” can be alleviated through the use of dimension reducing techniques such as feature (gene) selection which remove unnecessary features. There are many versions of feature selection, with varying biases and predictive abilities. However, predictive power is but one factor to consider when choosing a feature selection technique: one must also consider the technique's stability, that is, its ability to create feature subsets which remain valid in the face of changes to the data. While there has been work in determining the relative stability of different feature selection techniques, this does not always help determine whether a chosen feature selection technique will give stable feature subsets for a specific dataset. Factors such as difficulty of learning (e.g., dataset difficulty) may also influence feature selection stability, making generally-true facts about different techniques not applicable to a given dataset. In this work, we study how dataset difficulty can affect the stability of feature selection techniques, leading to good performance from bad techniques and vice versa. We use a set of twenty-six DNA microarray datasets with varying levels of difficulty of learning, along with four levels of dataset perturbation, six feature selection techniques with various levels of stability, and twelve feature subset sizes. The results show that as the dataset difficulty increases, the stability decreases. However, the relative stability between the techniques remains the same. Additionally, the more difficult the dataset, the more the stability is affected by changes to the data. We also found that unstable rankers are more affected by the transition between Easy and Moderate datasets, whereas the stable techniques are more affected by the change between Moderate and Hard datasets. Lastly, as the feature subset size increases, the stability increases and the difference between the levels of dataset difficulty decreases. Overall, we conclude that difficulty of learning must be taken into account before interpreting stability results.

Similar Papers
  • Book Chapter
  • Citations664

Robust Feature Selection Using Ensemble Feature Selection Techniques

  • Jan 01, 2008
  • Yvan Saeys +2
  • Conference Article
  • Citations2

Determining the Number of Iterations Appropriate for Ensemble Gene Selection on Microarray Data

  • Dec 01, 2012
  • David J Dittman +3
  • PDF
  • Research Article
  • Citations1

Feature Selection Techniques and Classification Accuracy of Supervised Machine Learning in Text Mining

  • May 01, 2019
  • Journal of Information Engineering and Applications
  • Loise Makara +2
  • Conference Article
  • Citations15

Comparison of Stability for Different Families of Filter-Based and Wrapper-Based Feature Selection

  • Dec 01, 2013
  • Randall Wald +2
  • Dissertation

Evolutionary Computation for Feature Selection in Classification

  • Jan 01, 2018
  • Hoai Nguyen
  • Book Chapter
  • Citations2

A Comparative Study on Feature Selection Techniques for Multi-cluster Text Data

  • Aug 24, 2018
  • Ananya Gupta +1
  • Research Article
  • Citations55

Feature selection with limited datasets

  • Oct 01, 1999
  • Medical Physics
  • Matthew A Kupinski +1
  • Research Article

Enhancing COVID-19 Prediction Using Machine Learning: A Comparative Analysis of Feature Selection and Classification Techniques

  • Mar 24, 2025
  • Journal of Information Systems Engineering and Management
  • L William Mary
  • Conference Article

Credit Risk Assessment: A Comparative Study of Feature Selection and Oversampling Techniques on High-Dimensional Imbalanced Datasets

  • Aug 28, 2025
  • Sudhansu Ranjan Lenka +4
  • Research Article
  • Citations15

Feature Selection Techniques on Thyroid, Hepatitis, and Breast Cancer Datasets

  • Mar 31, 2013
  • International Journal on Data Mining and Intelligent Information Technology Applications
  • Mohammad Ashraf - +2
  • Dissertation
  • Citations4

Particle Swarm Optimisation for Feature Selection in Classification

  • Jan 01, 2014
  • Bing Xue
  • Preprint Article
  • Citations5

A Sparse-Modeling based approach for Class-Specific feature selection

  • May 17, 2019
  • Davide Nardone +2
  • Conference Article
  • Citations5

Video classification based on ConvNet collaboration and feature selection

  • May 01, 2017
  • Emel Boyaci +1
  • Conference Article
  • Citations7

Effects of the Use of Boosting on Classification Performance of Imbalanced Bioinformatics Datasets

  • Nov 01, 2014
  • Taghi M Khoshgoftaar +3
  • Conference Article
  • Citations3

Ensemble of Filter and Embedded Feature Selection Techniques for Malware Classification using High-dimensional Jar Extension Dataset

  • Feb 23, 2023
  • Yi Wei Tye +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.