• Home
  • Search
  • Does Choice in Model Selection Affect Maximum Likelihood Analysis?
  • Cite Icon148
  • https://doi.org/10.1080/10635150801898920Copy DOI Icon

Does Choice in Model Selection Affect Maximum Likelihood Analysis?

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In order to have confidence in model-based phylogenetic analysis, the model of nucleotide substitution adopted must be selected in a statistically rigorous manner. Several model-selection methods are applicable to maximum likelihood (ML) analysis, including the hierarchical likelihood-ratio test (hLRT), Akaike information criterion (AIC), Bayesian information criterion (BIC), and decision theory (DT), but their performance relative to empirical data has not been investigated thoroughly. In this study, we use 250 phylogenetic data sets obtained from TreeBASE to examine the effects that choice in model selection has on ML estimation of phylogeny, with an emphasis on optimal topology, bootstrap support, and hypothesis testing. We show that the use of different methods leads to the selection of two or more models for approximately 80% of the data sets and that the AIC typically selects more complex models than alternative approaches. Although ML estimation with different best-fit models results in incongruent tree topologies approximately 50% of the time, these differences are primarily attributable to alternative resolutions of poorly supported nodes. Furthermore, topologies and bootstrap values estimated with ML using alternative statistically supported models are more similar to each other than to topologies and bootstrap values estimated with ML under the Kimura two-parameter (K2P) model or maximum parsimony (MP). In addition, Swofford-Olsen-Waddell-Hillis (SOWH) tests indicate that ML trees estimated with alternative best-fit models are usually not significantly different from each other when evaluated with the same model. However, ML trees estimated with statistically supported models are often significantly suboptimal to ML trees made with the K2P model when both are evaluated with K2P, indicating that not all models perform in an equivalent manner. Nevertheless, the use of alternative statistically supported models generally does not affect tests of monophyletic relationships under either the Shimodaira-Hasegawa (S-H) or SOWH methods. Our results suggest that although choice in model selection has a strong impact on optimal tree topology, it rarely affects evolutionary inferences drawn from the data because differences are mainly confined to poorly supported nodes. Moreover, since ML with alternative best-fit models tends to produce more similar estimates of phylogeny than ML under the K2P model or MP, the use of any statistically based model-selection method is vastly preferable to forgoing the model-selection process altogether.

Similar Papers
  • Research Article

Evaluating Life Time Models: A Comparative Study of the Inverse Weibull and Inverse Lognormal Distributions

  • Jan 01, 2026
  • Journal of Mathematical Sciences & Computational Mathematics
  • Research Article
  • Citations5

Evaluation on genetic relationships among China’s endemic Curcuma L. herbs by mtDNA

  • Jan 01, 2018
  • Phyton
  • Deng Jb +6
  • Research Article
  • Citations56

Selecting the model for multiple imputation of missing data: Just use an IC!

  • Feb 24, 2021
  • Statistics in Medicine
  • Firouzeh Noghrehchi +3
  • Research Article
  • Citations58

PREDICTION/ESTIMATION WITH SIMPLE LINEAR MODELS: IS IT REALLY THAT SIMPLE?

  • Dec 06, 2006
  • Econometric Theory
  • Yuhong Yang
  • Research Article

First Report of powdery mildew of Quercus guyavifolia (Fagaceae) Caused by Erysiphe quercicola.

  • May 28, 2024
  • Plant Disease
  • Yi Zhang +6
  • Research Article
  • Citations638

Maximum Likelihood and Minimum-Steps Methods for Estimating Evolutionary Trees from Data on Discrete Characters

  • Sep 01, 1973
  • Systematic Biology
  • J Felsenstein
  • Research Article
  • Citations1

Comparative Study of Blu and ML Estimators of Location and Scale Parameters of an Extreme Value Distribution for Small Samples and Type I Censorship

  • Jan 01, 1989
  • Communications in Statistics - Simulation and Computation
  • Mohamed M Bugaighis
  • PDF
  • Research Article

Modelling the Botswana Pula/Us Dollar exchange rate using the Skewed generalized t (SGT) distributions

  • Dec 21, 2021
  • International Journal of Scientific Research and Management
  • Wilson Moseki Thupeng
  • Research Article
  • Citations2

Improved inference for MCP-Mod approach using time-to-event endpoints with small sample sizes.

  • Apr 29, 2023
  • Pharmaceutical statistics
  • Márcio A Diniz +2
  • Research Article
  • Citations6

Draft Genome Sequences Resources of Mulberry Dwarf Phytoplasma Strain MDGZ-01 Associated with Mulberry Yellow Dwarf (MYD) Diseases.

  • Jun 21, 2022
  • Plant disease
  • Longhui Luo +5
  • Research Article
  • Citations4

Selecting the Best Growth Model for Fish Using Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC)

  • Sep 01, 2010
  • 臺灣水產學會刊
  • Sher Khan Panhwar +2
  • Research Article
  • Citations2

Restricted maximum likelihood estimation under Eisenhart model Ill

  • Sep 01, 1991
  • Statistica Neerlandica
  • K.R Lee +1
  • Research Article
  • Citations94

Federated Edge Learning With Misaligned Over-the-Air Computation

  • Jun 01, 2022
  • IEEE Transactions on Wireless Communications
  • Yulin Shao +2
  • Research Article
  • Citations25

MONOPHYLY OF THE GENUS CLOSTERIUM AND THE ORDER DESMIDIALES (CHAROPHYCEAE, CHLOROPHYTA) INFERRED FROM NUCLEAR SMALL SUBUNIT rDNA DATA

  • Dec 01, 2001
  • Journal of Phycology
  • Takashi Denboh +2
  • Research Article
  • Citations197

A review and comparison of four commonly used Bayesian and maximum likelihood model selection tools

  • Nov 26, 2007
  • Ecological Modelling
  • Eric J Ward
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.