- Research Article
- 10.1097/01720610-200804000-00006
Considerations of sample size in medical research
- Apr 01, 2008
- Journal of the American Academy of Physician Assistants
- John W Waterbor + 1 more +1
Considerations of sample size in medical research
Several methods of factor extraction have recently gained popularity as a procedure for dealing with estimation problems associated with small sample sizes, which can be found in the various behavioral science disciplines, such as comparative psychology and behavior genetics. Two popular approaches for particularly small samples (below 50) include unweighted least squares factor analysis (ULS-FA) and regularized exploratory factor analysis (REFA). However, it is unclear how well each of the approaches performs with small samples in the context of exploratory bifactor modeling. In the current study, a comprehensive simulation study was conducted to evaluate the small sample behavior of the two approaches in terms of bifactor structure recovery under different sample size, factor loading, number of variables per factor, number of factors, and factor correlation experimental conditions. The results show that REFA is recommended for use over ULS-FA, particularly in the conditions involving low factor loadings, few group factors, or a small number of variables per factor.
Loading PDF
Considerations of sample size in medical research
Considerations of sample size in medical research
Assessing the impact of restricted follow-up and small sample sizes on survival estimations in prostate cancer using registry data.
294 Background: Economic evaluations in oncology aim to assess the value of new therapies in the long term based on clinical trial data that often have restricted follow-up times (< 5 years) and small sample sizes (< 500 patients). This requires the use of extrapolation assumptions on long-term survival that go beyond the observed data. In this analysis, differences in survival extrapolation methods are tested in samples of sizes and follow-up reflecting typical clinical trials against a background of known survival in prostate cancer from a US based cancer registry. Methods: Data from the National Cancer Institute's Surveillance Epidemiology and End Results (SEER) registry on long-term survival in patients with stage IV prostate cancer were employed. The data set comprised those patients diagnosed between 1988 and 2003, with follow-up data available until 2012. Additional survival for those who received surgery (compared to those who did not), was estimated based on extrapolations using standard parametric statistical models (exponential, Weibull, log-logistic, log-normal, Gamma) fitted to the observed data. Survival analyses were run for 5 sample size scenarios (n = 27,670, 1000, 500, 200, 50) and 6 follow-up scenarios (follow-up years = 25, 20, 10, 5, 2, 1) yielding 30 combination scenarios. Performance of the methods was tested relative to the maximum follow-up, maximum sample size scenario (i.e. reference case) from the SEER registry. Results: Log-logistic and log-normal models were associated with flat tails which led to inflated survival estimations. For scenarios with smaller sizes, gamma models often did not converge. Exponential models were the most frequently reported as best model fit (in approximately 50% of scenarios). Also, gains in OS were consistent when exponential models were selected, and closely matched gain in OS from the reference case. Conclusions: Since clinical trials in oncology are often associated with small patient sample sizes and restricted follow-up, selecting an exponential model may lead to the most consistent and stable results based on the experiment constructed here. Further research should confirm these results for other types of cancer.
Read moreSimilarity-Principle-Based Machine Learning Method for Clinical Trials and Beyond
With recent success in supervised learning, artificial intelligence (AI) and machine learning (ML) can play a vital role in precision medicine. Deep learning neural networks have been used in drug discovery when larger data is available. However, applications of machine learning in clinical trials with small sample size (around a few hundreds) are limited. We propose a Similarity-Principle-Based Machine Learning (SBML) method, which is applicable for small and large sample size problems. In SBML, the attribute-scaling factors are introduced to objectively determine the relative importance of each attribute (predictor). The gradient method is used in learning (training), that is, updating the attribute-scaling factors. We evaluate SBML when the sample size is small and investigate the effects of tuning parameters. Simulations show that SBML achieves better predictions in terms of mean squared errors for various complicated nonlinear situations than full linear models, optimal and ridge regressions, mixed effect models, support vector machine and decision tree methods.
Read moreAutoregressive Prediction with Rolling Mechanism for Time Series Forecasting with Small Sample Size
Reasonable prediction makes significant practical sense to stochastic and unstable time series analysis with small or limited sample size. Motivated by the rolling idea in grey theory and the practical relevance of very short‐term forecasting or 1‐step‐ahead prediction, a novel autoregressive (AR) prediction approach with rolling mechanism is proposed. In the modeling procedure, a new developed AR equation, which can be used to model nonstationary time series, is constructed in each prediction step. Meanwhile, the data window, for the next step ahead forecasting, rolls on by adding the most recent derived prediction result while deleting the first value of the former used sample data set. This rolling mechanism is an efficient technique for its advantages of improved forecasting accuracy, applicability in the case of limited and unstable data situations, and requirement of little computational effort. The general performance, influence of sample size, nonlinearity dynamic mechanism, and significance of the observed trends, as well as innovation variance, are illustrated and verified with Monte Carlo simulations. The proposed methodology is then applied to several practical data sets, including multiple building settlement sequences and two economic series.
Read moreFrequent Diagnostic Under-Grading in Isocitrate Dehydrogenase Wild-Type Gliomas due to Small Pathological Tissue Samples.
In contrast to isocitrate dehydrogenase (IDH) mutation analysis, which is homogenous within a given tumor, diagnostic errors in histological analysis following the 2016 World Health Organization (WHO) classification could be due to small samples because of histological heterogeneity. To assess whether the sample size sent to histopathology influences the tumor grading in IDH wild-type gliomas. Histologically diagnosed WHO grade, sample volume, and preoperative tumor volume data of 111 patients aged who received resection of IDHwt gliomas between January 2007 and December 2015 at our hospital were evaluated. The differences between absolute and relative pathological sample sizes stratified by WHO grade were conducted using One-Way-Permutation-Test. With a mean sample size of 10.9 cc, 83.8% of patients were histologically diagnosed as WHO grade IV, while 16.2% of patients with a mean sample size of 2.62 cc were diagnosed as WHO grade II/III. One-Way-Permutation-Test showed a significant difference between absolute tissue samples stratified by WHO grade (P=.0374). The distribution of preoperative tumor volumes with WHO grade IV vs WHO grade II/III showed no significant difference (P=.8587). Of all tumors with a sample size>10 cc 100% were pathologically diagnosed as WHO grade IV and those with sample size>5 cc 93.5% were diagnosed as WHO grade IV. Small sample sizes are associated with a higher risk of under-estimating malignancy in histological grading in IDHwt gliomas. This study suggests a standard minimum sample size (>5cc) in every resection. Modalities of adjuvant treatment for IDHwt, WHO grade II/III gliomas need to reflect a prognosis that is only marginally better than of a glioblastoma.
Read moreMapping evergreen broad-leaved species spatial cover in Italian forests from Sentinel-2 time series using Deep Learning
Climate change-induced shifts, such as prolonged growing seasons and milder winters, coupled with land-use alterations like forest management abandonment, are reshaping species composition across European forests. A significant spread of evergreen broad-leaved species (EVEs) was observed in southern European forests, driven by global change dynamics. However, large-scale spatio-temporal analysis of these changes are lacking and emphasizes the necessity of mapping these dynamics. The TRACEVE projects (&#8220;Tracing the evergreen broad-leaved species and their spread&#8221;) primary goal is therefore&#160;tracking&#160; EVEs' cover, spread and diversity on a national-scale in Italy. As part of the project, this study focuses on satellite remote sensing, Deep Learning, and forest mapping, aiming at creating seamless maps quantifying the current degree of EVEs cover within forests in Italy.Challenges arise in the transitional zones between evergreen and deciduous forests, where EVEs initially spread in the understory of a deciduous canopy. Tracking EVEs at the edge of their range, where abundance is rare and in mixed forests is therefore&#160;difficult. Leveraging Sentinel-2 remote sensing time series covering the full annual phenological cycle addresses this challenge, utilizing leaf-on and leaf-off canopy conditions. Values of species cover derived from ad-hoc&#160;forest plot observations in Italian protected areas&#160;across a latitudinal gradient serve as initial training data, although the small sampling size (~1000 plots) poses challenges for the generalizability of time series extrinsic regression models, particularly when employing state-of-the-art Deep Learning architectures.The main aim of the study is therefore the development of a robust, Sentinel-2 based mapping procedure to track EVEs within Italian forests on a national-scale. A remote sensing time series extrinsic regression model based on a Deep Learning architecture for cover degree mapping will be developed. Sentinel-2 annual time series, along with derived indices&#160;serve as predictors for the models, while target cover will be derived from the plot observations. The study is built around exploring selected strategies&#160;to address the issue of large-scale model generalizability, given a small training sample size, that is recorded within small representative areas scattered across Italy.&#160;This entails assessing the efficacy of self-supervised pretraining methodologies applied to remote sensing time series within forested regions, pretraining on a more extensive forest database, and evaluating the viability of training data augmentation techniques.To validate and evaluate results on a national scale, an independent forest vegetation database containing around 17,000 forest plots sampled across Italy is employed. This extensive dataset enhances the understanding of EVEs' distribution within Italy's diverse forest ecosystems and can further enhance understanding of complex model results.In conclusion, the study combines advanced satellite remote sensing technologies, Deep Learning methodologies coupled with vegetation plot datasets to map current distribution of EVEs in Italian forests. The findings contribute to the TRACEVE project's future objectives but also offer insights into the challenges and opportunities of Deep Learning models in large-scale forest mapping applications.AcknowledgementsThis research has been conducted within the project &#8220;TRACEVE - Tracing the evergreen broad-leaved species and their spread&#8221; (I 6452-B) funded by the Austrian Science Fund (FWF).
Read moreA comparative study on model selection and multiple model fusion
There exist quite a few criteria for penalty-based model selection. Although they have various justifications for large sample problems, their performance under small or moderate sample size is unclear which hinders the development of model combination methods using the appropriate penalty term. In this paper, we assess the performance of seven model selection criteria based on linear regression models with unknown noise variance. We set the true data generation mechanism to be within the model set as well as outside the model set. In the latter case, soft model selection through multiple model fusion is proposed and its difference from Bayesian model averaging is highlighted. The penalty term used in each model selection criterion provides a natural link to estimate the model probability without assuming any prior knowledge of the unknown parameter. An important question is whether the estimated model probabilities are consistent when multiple models are fused for prediction or interpolation. We argue that strong consistency only holds under large sample regime while soft model selection can still be better than choosing a single model with small sample size. Our numerical results using different model selection criteria for polynomial fitting indicate that the conditional model estimator (CME) has the best performance in selecting the correct model order and fusing multiple models for prediction and interpolation. The minimum description length (MDL) based criteria are next to CME and outperform Bayesian information criterion (BIC) and Akaike information criterion (AIC) significantly.
Read moreExploratory Factor Analysis With Small Samples and Missing Data
ABSTRACTExploratory factor analysis (EFA) is an extremely popular method for determining the underlying factor structure for a set of variables. Due to its exploratory nature, EFA is notorious for being conducted with small sample sizes, and recent reviews of psychological research have reported that between 40% and 60% of applied studies have 200 or fewer observations. Recent methodological studies have addressed small size requirements for EFA models; however, these models have only considered complete data, which are the exception rather than the rule in psychology. Furthermore, the extant literature on missing data techniques with small samples is scant, and nearly all existing studies focus on topics that are not of primary interest to EFA models. Therefore, this article presents a simulation to assess the performance of various missing data techniques for EFA models with both small samples and missing data. Results show that deletion methods do not extract the proper number of factors and estimate the factor loadings with severe bias, even when data are missing completely at random. Predictive mean matching is the best method overall when considering extracting the correct number of factors and estimating factor loadings without bias, although 2-stage estimation was a close second.
Read moreSnack product consumer surveys: large versus small samples
Snack product consumer surveys: large versus small samples
Sample size calculations for single-arm survival studies using transformations of the Kaplan-Meier estimator.
In single-arm clinical trials with survival outcomes, the Kaplan-Meier estimator and its confidence interval are widely used to assess survival probability and median survival time. Since the asymptotic normality of the Kaplan-Meier estimator is a common result, the sample size calculation methods have not been studied in depth. An existing sample size calculation method is founded on the asymptotic normality of the Kaplan-Meier estimator using the log transformation. However, the small sample properties of the log transformed estimator are quite poor in small sample sizes (which are typical situations in single-arm trials), and the existing method uses an inappropriate standard normal approximation to calculate sample sizes. These issues can seriously influence the accuracy of results. In this paper, we propose alternative methods to determine sample sizes based on a valid standard normal approximation with several transformations that may give an accurate normal approximation even with small sample sizes. In numerical evaluations via simulations, some of the proposed methods provided more accurate results, and the empirical power of the proposed method with the arcsine square-root transformation tended to be closer to a prescribed power than the other transformations. These results were supported when methods were applied to data from three clinical trials.
Read moreComparison of Data Mining Classification Algorithms on Educational Data under Different Conditions
The purpose of this study was to examine the performance of Naive Bayes, k-nearest neighborhood, neural networks, and logistic regression analysis in terms of sample size and test data rate in classifying students according to their mathematics performance. The target population was 62728 students in the 15-year-old group who were participated in the Programme for International Student Assessment (PISA) in 2012 from The Organisation for Economic Co-operation and Development (OECD) countries. The performance of each algorithm was tested by using 11%, 22%, 33%, 44% and 55% of each dataset for small (500 students), medium (1000 students) and large (5000 students) sample sizes. 100 replications were performed for each analysis. As the evaluation criteria, accuracy rates, RMSE values, and total elapsed time were used. RMSE values for each algorithm were statistically compared by using Friedman and Wilcoxon tests. The results revealed that while the classification performance of the methods increased as the sample size increased, the increase of training data ratio had different effects on the performance of the algorithms. The Naive Bayes showed high performance even in small samples, performed the analyzes very quickly, and was not affected by the change in the training data ratio. Logistic regression analysis was the most effective method in large samples but had a poor performance in small samples. While neural networks showed a similar tendency, its overall performance was lower than Naive Bayes and logistic regression. The lowest performances in all conditions were obtained by the k-nearest neighborhood algorithm.
Read moreTesting Differential Item Functioning in Small Samples
Differential item functioning (DIF) is a pernicious statistical issue that can mask true group differences on a target latent construct. A considerable amount of research has focused on evaluating methods for testing DIF, such as using likelihood ratio tests in item response theory (IRT). Most of this research has focused on the asymptotic properties of DIF testing, in part because many latent variable methods require large samples to obtain stable parameter estimates. Much less research has evaluated these methods in small sample sizes despite the fact that many social and behavioral scientists frequently encounter small samples in practice. In this article, we examine the extent to which model complexity—the number of model parameters estimated simultaneously—affects the recovery of DIF in small samples. We compare three models that vary in complexity: logistic regression with sum scores, the 1-parameter logistic IRT model, and the 2-parameter logistic IRT model. We expected that logistic regression with sum scores and the 1-parameter logistic IRT model would more accurately estimate DIF because these models yielded more stable estimates despite being misspecified. Indeed, a simulation study and empirical example of adolescent substance use show that, even when data are generated from / assumed to be a 2-parameter logistic IRT, using parsimonious models in small samples leads to more powerful tests of DIF while adequately controlling for Type I error. We also provide evidence for minimum sample sizes needed to detect DIF, and we evaluate whether applying corrections for multiple testing is advisable. Finally, we provide recommendations for applied researchers who conduct DIF analyses in small samples.
Read moreDesign and analysis of a 2-year parallel follow-up of repeated ivermectin mass drug administrations for control of malaria: Small sample considerations for cluster-randomized trials with count data.
Cluster-randomized trials allow for the evaluation of a community-level or group-/cluster-level intervention. For studies that require a cluster-randomized trial design to evaluate cluster-level interventions aimed at controlling vector-borne diseases, it may be difficult to assess a large number of clusters while performing the additional work needed to monitor participants, vectors, and environmental factors associated with the disease. One such example of a cluster-randomized trial with few clusters was the "efficacy and risk of harms of repeated ivermectin mass drug administrations for control of malaria" trial. Although previous work has provided recommendations for analyzing trials like repeated ivermectin mass drug administrations for control of malaria, additional evaluation of the multiple approaches for analysis is needed for study designs with count outcomes. Using a simulation study, we applied three analysis frameworks to three cluster-randomized trial designs (single-year, 2-year parallel, and 2-year crossover) in the context of a 2-year parallel follow-up of repeated ivermectin mass drug administrations for control of malaria. Mixed-effects models, generalized estimating equations, and cluster-level analyses were evaluated. Additional 2-year parallel designs with different numbers of clusters and different cluster correlations were also explored. Mixed-effects models with a small sample correction and unweighted cluster-level summaries yielded both high power and control of the Type I error rate. Generalized estimating equation approaches that utilized small sample corrections controlled the Type I error rate but did not confer greater power when compared to a mixed model approach with small sample correction. The crossover design generally yielded higher power relative to the parallel equivalent. Differences in power between analysis methods became less pronounced as the number of clusters increased. The strength of within-cluster correlation impacted the relative differences in power. Regardless of study design, cluster-level analyses as well as individual-level analyses like mixed-effects models or generalized estimating equations with small sample size corrections can both provide reliable results in small cluster settings. For 2-year parallel follow-up of repeated ivermectin mass drug administrations for control of malaria, we recommend a mixed-effects model with a pseudo-likelihood approximation method and Kenward-Roger correction. Similarly designed studies with small sample sizes and count outcomes should consider adjustments for small sample sizes when using a mixed-effects model or generalized estimating equation for analysis. Although the 2-year parallel follow-up of repeated ivermectin mass drug administrations for control of malaria is already underway as a parallel trial, applying the simulation parameters to a crossover design yielded improved power, suggesting that crossover designs may be valuable in settings where the number of available clusters is limited. Finally, the sensitivity of the analysis approach to the strength of within-cluster correlation should be carefully considered when selecting the primary analysis for a cluster-randomized trial.
Read moreParasite prevalence and sample size: misconceptions and solutions
Parasite prevalence and sample size: misconceptions and solutions
The factor analysis procedure for exploration: a short guide with examples / El análisis factorial exploratorio: una guía breve con ejemplos
Surveys and tests contain multiple test items, sets of repeated tests or multiple survey questions. Commonly, these units are arranged within instruments subject to varying contexts of tests or questions. The analyst’s goal is to discover communalities across these items such that items can be reduced down to common meaningful factors. We provide a literature review that supports our further choice between exploratory analytical models for analysing empirical data and for building a guide to interpreting results. The purpose of this research is to provide a methodological and systematic framework for researchers who consider exploratory analyses. Following a comparison between factor extraction methods, we suggest various approaches to look at the association between the original variables and the factor, as well as correlations between factors. Our empirical case study data is a survey instrument of 19 items from a questionnaire developed by the Branco Weiss Institute in Israel, for evaluating at-risk high school and intermediate school students. Properties of the data such as the sample size, the quality of data by means of distribution patterns and extreme values, and correlations between the original items are considered. We argue that a concurrent integration of two fundamental processes — the empirical model fit and the substantive meaning — are essential in the process of implementing exploratory analysis results. The main conclusion is that the process of exploring latent factors needs an allocation of analytical resources, similar to other statistical modelling practices. Data and context mutually function as the platform for arriving at the optimal number of factors and their item composition. The exploratory factor analysis is a powerful tool for researchers who are ready to operate this tool properly.
Read more