• Home
  • Search
  • How Many Genes are Needed for a Discriminant Microarray Data Analysis
  • Open Access IconOpen Access
  • Cite Icon122
  • https://doi.org/10.1007/978-1-4615-0873-1_11Copy DOI Icon

How Many Genes are Needed for a Discriminant Microarray Data Analysis

  • Apr 6, 2001
  • Wentian Li +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The analysis of the leukemia data from Whitehead/MIT group is a discriminant analysis (also called a supervised learning). Among thousands of genes whose expression levels are measured, not all are needed for discriminant analysis: a gene may either not contribute to the separation of two types of tissues/cancers, or it may be redundant because it is highly correlated with other genes. There are two theoretical frameworks in which variable selection (or gene selection in our case) can be addressed. The first is model selection, and the second is model averaging. We have carried out model selection using Akaike information criterion and Bayesian information criterion with logistic regression (discrimination, prediction, or classification) to determine the number of genes that provide the best model. These model selection criteria set upper limits of 22∼25 and 12∼13 genes for this data set with 38 samples, and the best model consists of only one (no.4847, zyxin) or two genes. We have also carried out model averaging over the best single-gene logistic predictors using three different weights: maximized likelihood, prediction rate on training set, and equal weight. We have observed that the performance of most of these weighted predictors on the testing set is gradually reduced as more genes are included, but a clear cutoff that separates good and bad prediction performance is not found.Key wordsmodel/variable selectionAkaike/Bayesian information criterionmaximum likelihoodlogistic regressiondiscriminant analysis

Similar Papers
  • Research Article
  • Citations204

Adaptive Model Selection

  • Mar 01, 2002
  • Journal of the American Statistical Association
  • Xiaotong Shen +1
  • Research Article
  • Citations72

Latent profile analysis with nonnormal mixtures: A Monte Carlo examination of model selection using fit indices

  • Mar 10, 2015
  • Computational Statistics & Data Analysis
  • Grant B Morgan +2
  • Research Article

Evaluating Life Time Models: A Comparative Study of the Inverse Weibull and Inverse Lognormal Distributions

  • Jan 01, 2026
  • Journal of Mathematical Sciences & Computational Mathematics
  • Research Article
  • Citations2

Causal analysis of futures sugar prices in Zhengzhou

  • Dec 13, 2011
  • Acta Mathematicae Applicatae Sinica, English Series
  • Fang Wang +3
  • Research Article
  • Citations58

PREDICTION/ESTIMATION WITH SIMPLE LINEAR MODELS: IS IT REALLY THAT SIMPLE?

  • Dec 06, 2006
  • Econometric Theory
  • Yuhong Yang
  • Research Article
  • Citations89

Information criteria for model selection

  • Feb 20, 2023
  • WIREs Computational Statistics
  • Jiawei Zhang +2
  • Research Article
  • Citations24

Model Selection Strategies and the Use of Association Models to Detect Group Differences

  • May 01, 1994
  • Sociological Methods & Research
  • Raymond Sin-Kwok Wong
  • Research Article
  • Citations24

Variable Selection in Multivariable Regression UsingSAS/IML

  • Jan 01, 2002
  • Journal of Statistical Software
  • Ali A Al-Subaihi
  • Research Article
  • Citations16

Investigation on the Improvement of Prediction by Bootstrap Model Averaging

  • Jan 01, 2006
  • Methods of Information in Medicine
  • N H Augustin +2
  • Research Article
  • Citations14

False Discovery Rate (FDR) and Familywise Error Rate (FER) Rules for Model Selection in Signal Processing Applications

  • Jan 01, 2022
  • IEEE Open Journal of Signal Processing
  • Petre Stoica +1
  • Research Article
  • Citations12

A jackknife type approach to statistical model selection

  • Jul 29, 2011
  • Journal of Statistical Planning and Inference
  • Hyunsook Lee +2
  • Research Article
  • Citations37

Bayesian cross-validation for model evaluation and selection, withapplication to the North American Breeding Bird Survey.

  • Jul 01, 2016
  • Ecology
  • William A Link +1
  • Research Article
  • Citations10

Application of time series methods for dengue cases in North India (Chandigarh)

  • Nov 15, 2019
  • Journal of Public Health
  • Kumar Shashvat +2
  • Research Article
  • Citations56

Selecting the model for multiple imputation of missing data: Just use an IC!

  • Feb 24, 2021
  • Statistics in Medicine
  • Firouzeh Noghrehchi +3
  • Research Article
  • Citations148

Does Choice in Model Selection Affect Maximum Likelihood Analysis?

  • Feb 01, 2008
  • Systematic Biology
  • Jennifer Ripplinger +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.