- Research Article
- 10.14219/jada.archive.1933.0299
Mottled Enamel
- Oct 01, 1933
- The Journal of the American Dental Association
- J Scott Walker
Mottled Enamel
Bootstrap method for minimum message length autoregressive model order selection
Mottled Enamel
Mottled Enamel
A comparative study on model selection and multiple model fusion
There exist quite a few criteria for penalty-based model selection. Although they have various justifications for large sample problems, their performance under small or moderate sample size is unclear which hinders the development of model combination methods using the appropriate penalty term. In this paper, we assess the performance of seven model selection criteria based on linear regression models with unknown noise variance. We set the true data generation mechanism to be within the model set as well as outside the model set. In the latter case, soft model selection through multiple model fusion is proposed and its difference from Bayesian model averaging is highlighted. The penalty term used in each model selection criterion provides a natural link to estimate the model probability without assuming any prior knowledge of the unknown parameter. An important question is whether the estimated model probabilities are consistent when multiple models are fused for prediction or interpolation. We argue that strong consistency only holds under large sample regime while soft model selection can still be better than choosing a single model with small sample size. Our numerical results using different model selection criteria for polynomial fitting indicate that the conditional model estimator (CME) has the best performance in selecting the correct model order and fusing multiple models for prediction and interpolation. The minimum description length (MDL) based criteria are next to CME and outperform Bayesian information criterion (BIC) and Akaike information criterion (AIC) significantly.
Read moreModel Selection Strategies and the Use of Association Models to Detect Group Differences
This article discusses the use of association models to detect group differences and the choice of model selection criteria under different conditions. The performances of several commonly used model selection criteria in log-linear modeling are examined using Monte Carlo simulations. The results suggest that no single criterion can play the role of panacea in model selection. L2/ df, normed fit index (NFI) (or 1 - L12/L02), and Akaike's information criterion (AIC) are systematically biased toward models incorporating group differences. The log-likelihood ratio test (LRT), in conjunction with the nested chi-square difference test, is useful for samples with moderate sizes. Bayesian information criterion (BIC) is most reliable among the measures tested. It consistently provides correct information about group differences in association, especially when the sample size is large. Although BIC may give equivocal results under certain conditions, the ambiguities can be mitigated by an inclusion of competing models and a cautious interpretation of any small improvement in BIC.
Read moreModel selection for factor analysis: Some new criteria and performance comparisons
ABSTRACTThis paper derives Akaike information criterion (AIC), corrected AIC, the Bayesian information criterion (BIC) and Hannan and Quinn’s information criterion for approximate factor models assuming a large number of cross-sectional observations and studies the consistency properties of these information criteria. It also reports extensive simulation results comparing the performance of the extant and new procedures for the selection of the number of factors. The simulation results show the difficulty of determining which criterion performs best. In practice, it is advisable to consider several criteria at the same time, especially Hannan and Quinn’s information criterion, Bai and Ng’s ICp2 and BIC3, and Onatski’s and Ahn and Horenstein’s eigenvalue-based criteria. The model-selection criteria considered in this paper are also applied to Stock and Watson’s two macroeconomic data sets. The results differ considerably depending on the model-selection criterion in use, but evidence suggesting five factors for the first data and five to seven factors for the second data is obtainable.
Read moreVelocity Variables: Determining Predictive Metrics during the Back Squat and Bench Press to Failure at Different Relative Loads.
Lawson, DJ, Olmos, AA, Mosiman, SJ, Sontag, SA, Goodin, JR, and Dawes, JJ. Velocity Variables: Determining Predictive Metrics during the Back Squat and Bench Press to Failure at Different Relative Loads. J Strength Cond Res XX(X): 000-000, 2025-This study aimed to determine a best velocity variable and prediction model for estimating repetitions to failure (RTF) across 3 relative loads (%1RM) comparing the average concentric velocity (ACV) of a single repetition set (ACVSingle), ACV from the first repetition (ACVFirst), ACV across all repetitions (ACVMean), and the fastest repetition velocity (FRV) achieved during the back squat and bench press exercises. Twenty-six (n = 26; males = 18, females = 8) resistance-trained individuals performed 3 sets to failure at 90, 80, and 70% of their 1RM on both exercises for 2 testing sessions. Repeated measures mixed effects models were constructed for univariate, adjusted (corrected for sex), and interaction (velocity*sex) models from Visit 2 data. Model selection criteria were determined by the smallest residual mean error (RME) and standard deviation (SD), Akaike Information Criterion (AIC), and Bayesian Information Criterion (BIC) serving as fit indicators. Best fit models were cross-validated by applying fixed-effects coefficients from Visit 2 to Visit 3 velocity variables, estimating RTF and calculating error as the predicted versus observed variable delta. The ACVSingle adjusted model demonstrated the best fit for the squat (RME = 0.0056, SD = 3.7731, AIC = 360.88, BIC = 363.24). The FRV interaction model demonstrated the best fit for the bench press (RME = 0.0303, SD = 2.4011, AIC = 300.71, and BIC = 303.07). Although no single predictor exhibited superiority across all intensities, ACVSingle and FRV provide lower prediction error variability under specific conditions, with the best predictor determined by both intensity and exercise.
Read moreHybrid modeling approaches for predicting COVID-19 mortality: A comparative study across USA, France, and India
Hybrid modeling approaches for predicting COVID-19 mortality: A comparative study across USA, France, and India
Variable Selection in Multivariable Regression UsingSAS/IML
This paper introduces a SAS/IML program to select among the multivariate model candidates based on a few well-known multivariate model selection criteria. Stepwise regression and all-possible-regression are considered. The program is user friendly and requires the user to paste or read the data at the beginning of the module, include the names of the dependent and independent variables (the y's and the x's), and then run the module. The program produces the multivariate candidate models based on the following criteria: Forward Selection, Forward Stepwise Regression, Backward Elimination, Mean Square Error, Coefficient of Multiple Determination, Adjusted Coefficient of Multiple Determination, Akaike's Information Criterion, the Corrected Form of Akaike's Information Criterion, Hannan and Quinn Information Criterion, the Corrected Form of Hannan and Quinn (HQc) Information Criterion, Schwarz's Criterion, and Mallow's PC. The output also constitutes detailed as well as summarized results.
Read moreA review and comparison of four commonly used Bayesian and maximum likelihood model selection tools
A review and comparison of four commonly used Bayesian and maximum likelihood model selection tools
Bayesian model selection in ARFIMA models
Bayesian model selection in ARFIMA models
Evaluating Life Time Models: A Comparative Study of the Inverse Weibull and Inverse Lognormal Distributions
This study presents a comparative analysis of the Inverse Weibull and Inverse Lognormal distributions using both simulated and real-world data with Maximum Likelihood Estimation (MLE).Model selection criteria included Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Anderson-Darling (AD), and Kolmogorov-Smirnov (KS) tests.Six simulated sample sizes (50, 100, 150, 200, 250, and 500) were used, with 1000 replications each.Results showed the Inverse Lognormal distribution consistently had lower AIC and BIC values at small to moderate sample sizes.Furthermore, real-world stock price data (sample size 100) from the Nigerian Stock Exchange was analyzed.Descriptive statistics and goodness-of-fit tests favored the Inverse Lognormal model.These findings support the utility of AIC, BIC, AD, and KS in lifetime data modeling and model selection.
Read moreInformation criteria for model selection
The rapid development of modeling techniques has brought many opportunities for data‐driven discovery and prediction. However, this also leads to the challenge of selecting the most appropriate model for any particular data task. Information criteria, such as the Akaike information criterion (AIC) and Bayesian information criterion (BIC), have been developed as a general class of model selection methods with profound connections with foundational thoughts in statistics and information theory. Many perspectives and theoretical justifications have been developed to understand when and how to use information criteria, which often depend on particular data circumstances. This review article will revisit information criteria by summarizing their key concepts, evaluation metrics, fundamental properties, interconnections, recent advancements, and common misconceptions to enrich the understanding of model selection in general.This article is categorized under:Data: Types and Structure > Traditional Statistical DataStatistical Learning and Exploratory Methods of the Data Sciences > Modeling MethodsStatistical and Graphical Methods of Data Analysis > Information Theoretic MethodsStatistical Models > Model Selection
Read moreProcess-Monitoring-for-Quality — A Model Selection Criterion for l - Regularized Logistic Regression
Process-Monitoring-for-Quality — A Model Selection Criterion for l - Regularized Logistic Regression
Model comparison and selection for stationary space–time models
Model comparison and selection for stationary space–time models
New Criteria of Model Selection and Model Averaging in Linear Regression Models
Model selection is an important part of any statistical analysis. Many tools are suggested for selecting the best model including frequentist and Bayesian perspectives. There is often a considerable uncertainty in the selection of a particular model to be the best approximating model. Model selection uncertainty arises when the data are used for both model selection and parameter estimation. Bias in estimators of model parameters often arise when data based selection has been done. Therefore, model averaging of the parameter estimators will be done to alleviate the bias in model selection in a set of candidate models, by combining the information from a set of candidate models. This paper is two-fold, new criteria of model selection are proposed based on different averages of AIC, BIC, AICc, and HQC. Also, model averaging is introduced to compare the parameter estimators in model averaging with the ones in model selection. Two Simulation studies are considered, the first is for model selection and showed that the new proposed criteria are lies between some of the known criteria such as AIC, BIC, AICc, and HQC, and so they can be used as new criteria of model selection. The second simulation study is for model averaging and showed that the parameter estimators have less bias and less predicted mean square error (PMSE) compared with the parameter estimators in model selection.
Read moreLatent profile analysis with nonnormal mixtures: A Monte Carlo examination of model selection using fit indices
Latent profile analysis with nonnormal mixtures: A Monte Carlo examination of model selection using fit indices