- Research Article
- 10.1097/01720610-200804000-00006
Considerations of sample size in medical research
- Apr 01, 2008
- Journal of the American Academy of Physician Assistants
- John W Waterbor + 1 more +1
Considerations of sample size in medical research
Whereas general sample size guidelines have been suggested when estimating multilevel models, they are only generalizable to a relatively limited number of data conditions and model structures, both of which are not very feasible for the applied researcher. In an effort to expand our understanding of two-level multilevel models under less than ideal conditions, Monte Carlo methods, through SAS/IML, were used to examine model convergence rates, parameter point estimates (statistical bias), parameter interval estimates (confidence interval accuracy and precision), and both Type I error control and statistical power of tests associated with the fixed effects from linear two-level models estimated with PROC MIXED. These outcomes were analyzed as a function of: (a) level-1 sample size, (b) level-2 sample size, (c) intercept variance, (d) slope variance, (e) collinearity, and (f) model complexity. Bias was minimal across nearly all conditions simulated. The 95% confidence interval coverage and Type I error rate tended to be slightly conservative. The degree of statistical power was related to sample sizes and level of fixed effects; higher power was observed with larger sample sizes and level-1 fixed effects.
Considerations of sample size in medical research
Considerations of sample size in medical research
Conflated random slopes in multilevel analysis
Multilevel analysis is widely used for hierarchical data (e.g., individuals within clusters), allowing researchers to examine how individual- and cluster-level (i.e., level 1 and level 2) attributes relate to individual outcomes. A well-known issue concerns conflated fixed slopes, which arise when a level 1 predictor with systematic variance across clusters is modeled with a single fixed effect at the individual level, conflating its within- and between-cluster slopes. Less recognized, however, is the problem of conflated random slopes that occurs when the sources of heterogeneity introduced by the within- and between-cluster components of a level 1 predictor are not separated (i.e., the fixed effects of the predictor are separated across levels, but not their random effects). A conflated random slope is a blend of slope heterogeneity and intercept heteroscedasticity, producing misleading conclusions about between-cluster differences. Proper parameterization of the random slopes is necessary to avoid this issue, either using cluster-mean-centering (CMC) or constant-centering (CC). Despite its importance, the practical consequences of conflated random slopes remain rarely discussed in psychology and educational research. This study addresses that gap using a Monte Carlo simulation comparing three models that estimate unconflated random slopes and one widely used model that estimates conflated random slopes. The models differed in their parameterization of the random slope for a level 1 predictor: The within-only model had a random slope for the CMC level 1 predictor only; the within–between model had random slopes for both the CMC level 1 predictor and the CC level 2 predictor; the within–contextual model had random slopes for the CC level 1 predictor and the CC level 2 predictor; and the random-conflated model had a random slope for the CC level 1 predictor only. The simulation factors were level 1 sample size (5 and 30), level 2 sample size (25 and 50), random slope reliability (0.20 and 0.60), and the ratio of intercept heteroscedasticity to slope heterogeneity (0.5, 1.0, 1.5). After examining model convergence, model performance was assessed in terms of bias in fixed and random effects, power to detect random slopes, and accuracy of standard errors. Results showed clear trade-offs. The random-conflated model had very high convergence (> .97) even in small samples, but it consistently underestimated level 1 slope variance when within- and between-cluster variances differed, yielding downwardly biased variance estimates. The within-only model correctly specified the random slope but imposed homogeneity in intercept variance across cluster means, leading to slightly biased SEs for level 2 fixed effects. By contrast, the within–between and within–contextual models—theoretically correct specifications—had substantially lower convergence rates (.38–.73) in small samples but generally outperformed the simpler alternatives in parameter estimation and SE accuracy in larger sample sizes. Regarding power, all four models showed high capability to detect non-zero level 1 random slope variance. However, power to detect level 2 random slope variance in the within–between and within–contextual models was modest (.39 and .58, respectively), especially when sample sizes at level 1 and level 2 were small. This highlights the tension in practice between model complexity and the available information for the model from the data. Overall, our results show that estimating conflated random slopes distorts both fixed and random effect estimates, threatening valid inferences about between-cluster differences. When sample sizes permit, researchers may specify the within–between model, particularly if intercept heteroscedasticity matters. When data are limited or intercept heteroscedasticity is not of interest, the within-only model is a reasonable alternative to avoid random conflation, though with the trade-off of slightly biased SEs for level 2 fixed effects.
Read moreSample Size and Statistical Power Calculation in Genetic Association Studies
A sample size with sufficient statistical power is critical to the success of genetic association studies to detect causal genes of human complex diseases. Genome-wide association studies require much larger sample sizes to achieve an adequate statistical power. We estimated the statistical power with increasing numbers of markers analyzed and compared the sample sizes that were required in case-control studies and case-parent studies. We computed the effective sample size and statistical power using Genetic Power Calculator. An analysis using a larger number of markers requires a larger sample size. Testing a single-nucleotide polymorphism (SNP) marker requires 248 cases, while testing 500,000 SNPs and 1 million markers requires 1,206 cases and 1,255 cases, respectively, under the assumption of an odds ratio of 2, 5% disease prevalence, 5% minor allele frequency, complete linkage disequilibrium (LD), 1:1 case/control ratio, and a 5% error rate in an allelic test. Under a dominant model, a smaller sample size is required to achieve 80% power than other genetic models. We found that a much lower sample size was required with a strong effect size, common SNP, and increased LD. In addition, studying a common disease in a case-control study of a 1:4 case-control ratio is one way to achieve higher statistical power. We also found that case-parent studies require more samples than case-control studies. Although we have not covered all plausible cases in study design, the estimates of sample size and statistical power computed under various assumptions in this study may be useful to determine the sample size in designing a population-based genetic association study.
Read moreTwo regression methods for estimation of a two-parameter Weibull distribution for reliability of dental materials
Two regression methods for estimation of a two-parameter Weibull distribution for reliability of dental materials
Influence of group sample size on statistical power of tests for quantitative data with an imbalanced design
To explore the relationship between sample size in the groups and statistical power of ANOVA and Kruskal-Wallis H test with an imbalanced design. The sample sizes of the two tests were estimated by SAS program with given parameter settings, and Monte Carlo simulation was used to examine the changes in power when the total sample size varied or remained fixed. In ANOVA, when the total sample size was fixed, increasing the sample size in the group with a larger mean square error improved the statistical power, but an excessively large difference in the sample sizes between groups led to reduced power. When the total sample size was not fixed, a larger mean square error in the group with increased sample size was associated with a greater increase of the statistical power. In Kruskal-wallis H test, when the total sample size was fixed, increasing the sample size in groups with large mean square errors increased the statistical power irrespective of the sample size difference between the groups; when total sample size was not fixed, a larger mean square error in the group with increased sample size resulted in an increased statistical power, and the increment was similar to that for a fixed total sample size. The relationship between statistical power and sample size in groups is affected by the mean square error, and increasing the sample size in a group with a large mean square error increases the statistical power. In Kruskal-Wallis H test, increasing the sample size in a group with a large mean square error is more cost- effective than increasing the total sample size to improve the statistical power.
Read moreSample size and power calculations in Mendelian randomization with a single instrumental variable and a binary outcome
Background: Sample size calculations are an important tool for planning epidemiological studies. Large sample sizes are often required in Mendelian randomization investigations.Methods and results: Resources are provided for investigators to perform sample size and power calculations for Mendelian randomization with a binary outcome. We initially provide formulae for the continuous outcome case, and then analogous formulae for the binary outcome case. The formulae are valid for a single instrumental variable, which may be a single genetic variant or an allele score comprising multiple variants. Graphs are provided to give the required sample size for 80% power for given values of the causal effect of the risk factor on the outcome and of the squared correlation between the risk factor and instrumental variable. R code and an online calculator tool are made available for calculating the sample size needed for a chosen power level given these parameters, as well as the power given the chosen sample size and these parameters.Conclusions: The sample size required for a given power of Mendelian randomization investigation depends greatly on the proportion of variance in the risk factor explained by the instrumental variable. The inclusion of multiple variants into an allele score to explain more of the variance in the risk factor will improve power, however care must be taken not to introduce bias by the inclusion of invalid variants.
Read moreNeed for equivalence testing of efficacy of alternative antibiotics for treatment of pertussis.
Need for equivalence testing of efficacy of alternative antibiotics for treatment of pertussis.
Assessing Methods for Predictive Cut-Point Estimation: A Simulation-Based Comparison
IntroductionThe identification of an optimal cut-point for continuous biomarkers plays a crucial role in defining patient subgroups likely to benefit from specific treatments. While the literature has extensively covered prognostic biomarkers, those that provide outcome prediction regardless of treatment, the methodological framework for identifying predictive effect, which inform treatment effect heterogeneity, is less developed. This is primarily due to the added complexity of modelling treatment-biomarker interactions, which poses challenges related to statistical power, overfitting, and bias. ObjectivesThis study aimed to compare three statistical methods for the identification of predictive cut-points in time-to-event data. Our goal was to assess their performance in estimating the correct interaction effect and identifying a responder subgroup, under simulation settings that account for variability in treatment efficacy, biomarker predictive effect, and subgroup prevalence. MethodsWe implemented three approaches: Procedure B of the Biomarker-Adaptive Threshold Design (M1), which combines test statistics across possible cut-points using a permutation test based on likelihood-ratio statistics; the Differential Hazard Ratio method (M2), which selects the cut-point with the largest difference in HRs across adjacent thresholds; and a Minimum P-value method (M3) adapted for interaction terms in the Cox model [1,2]. We conducted a simulation study with 1000 replications from an exponential distribution with an expected censoring rate of approximately 40%. Eight main scenarios were defined by all possible combinations of two sample sizes (n = 300 and n = 500), two treatment effect sizes (HR = 1 or 0.5), two interaction effect sizes (HR = 1 or 0.5), and a biomarker prognostic effect set to HR = 0.6. In addition, we included two extra scenarios calibrated to achieve 80% power: one based on the interaction effect test (β for treatment-biomarker interaction) and one on the subgroup effect test (β within responders). In each replication, the true cut-point was randomly drawn from the biomarker distribution between the 20th and 80th percentiles. For each method, we evaluated statistical power, cut-point estimation bias, subgroup and predictive coefficient estimation bias, and type I error. A significance level of 0.05 was used for all three methods. The procedures were also evaluated on a real case on a prostate cancer clinical trial conducted by the Second Veterans Administration Cooperative Urologic Research Group [3]. ResultsM1 consistently demonstrated robust performance, with type I error close to the nominal level ( , 5.6%) and minimal bias in cut-point estimation ( , 0.005±0.06). It maintained good power even when the subgroup size was small. M2 showed unstable cut-point estimates ( , 0.055±0.42) and high variability in interaction estimates ( , 0.463±1.46), yielding a very low power ( , 16.2%). While the M3 achieved the highest power in some scenarios ( , 82.1%), it exhibited significant type I error inflation ( , 50.1%) and substantial bias due to multiple testing without correction ( , -0.401±1.730). In small subgroups, all methods experienced reduced performance, but M1 remained the most stable. On the prostate cancer dataset, M1 identified a plausible treatment-responsive subgroup, while the other two methods produced conflicting or less reliable results. Conclusions Our results highlight the need for robust methods in predictive cut-point estimation. M1 showed the best balance between error control and accuracy. In contrast, M2 and M3 may lead to overfitting, unstable estimates, and inflated first error rates. Future research should extend these comparisons to more complex models including multivariate biomarkers.
Read moreBiostatistics for gastroenterologists. Part II ̶ Rethinking sample size.
Sample size determination is a critical aspect of medical studies, influencing the reliability and generalizability of research findings. This article explores the importance of sample size in both basic and clinical research. The considerations for sample size differ based on the type of research, whether it is involving humans, animals, or cells. In basic research, a larger sample size is necessary to ensure statistical power and reliable results, enhancing the precision and generalizability of findings. In clinical research, determining an appropriate sample size is crucial for obtaining valid and clinically relevant results, ensuring sufficient statistical power to detect differences between treatment groups or validate intervention efficacy. Reporting sample size calculations accurately and complying with reporting guidelines, such as the CONSORT Statement, is essential for transparent and comprehensive research publications. Consulting a statistician is highly recommended to ensure appropriate sample size determination, enhance scientific rigor, and obtain reliable and clinically relevant findings in medical research.
Read moreA comprehensive review of group level model performance in the presence of heteroscedasticity: Can a single model control Type I errors in the presence of outliers?
A comprehensive review of group level model performance in the presence of heteroscedasticity: Can a single model control Type I errors in the presence of outliers?
Read morePower and sample size calculations: A review and computer program
Power and sample size calculations: A review and computer program
Sample size and statistical power in [15O]H2O studies of human cognition.
Determining the appropriate sample size is a crucial component of positron emission tomography (PET) studies. Power calculations, the traditional method for determining sample size, were developed for hypothesis-testing approaches to data analysis. This method for determining sample size is challenged by the complexities of PET data analysis: use of exploratory analysis strategies, search for multiple correlated nodes on interlinked networks, and analysis of large numbers of pixels that may have correlated values due to both anatomical and functional dependence. We examine the effects of variable sample size in a study of human memory, comparing large (n = 33), medium (n = 16,17), small (n = 11, 11, 11), and very small (n = 6,6,7,7,7) samples. Results from the large sample are assumed to be the "gold standard." The primary criterion for assessing sample size is replicability. This is evaluated using a hierarchically ordered group of parameters: pattern of peaks, location of peaks, number of peaks, size (volume) of peaks, and intensity of the associated t (or z) statistic. As sample size decreases, false negatives begin to appear, with some loss of pattern and peak detection; there is no corresponding increase in false positives. The results suggest that good replicability occurs with a sample size of 10-20 subjects in studies of human cognition that use paired subtraction comparisons of single experimental/baseline conditions with blood flow differences ranging from 4 to 13%.
Read moreResponse
Response
RNAseqPS: A Web Tool for Estimating Sample Size and Power for RNAseq Experiment
Sample size and power determination is the first step in the experimental design of a successful study. Sample size and power calculation is required for applications for National Institutes of Health (NIH) funding. Sample size and power calculation is well established for traditional biological studies such as mouse model, genome wide association study (GWAS), and microarray studies. Recent developments in high-throughput sequencing technology have allowed RNAseq to replace microarray as the technology of choice for high-throughput gene expression profiling. However, the sample size and power analysis of RNAseq technology is an underdeveloped area. Here, we present RNAseqPS, an advanced online RNAseq power and sample size calculation tool based on the Poisson and negative binomial distributions. RNAseqPS was built using the Shiny package in R. It provides an interactive graphical user interface that allows the users to easily conduct sample size and power analysis for RNAseq experimental design. RNAseqPS can be accessed directly at http://cqs.mc.vanderbilt.edu/shiny/RNAseqPS/.
Read moreLatent tree models for multivariate density estimation : algorithms and applications
Multivariate density estimation is a fundamental problem in Applied Statistics and Machine Learning. Given a collection of data sampled from an unknown distribution, the task is to approximately reconstruct the generative distribution. There are two different approaches to the problem, the parametric approach and the non-parametric approach. In the parametric approach, the approximate distribution is represented by a model from a predetermined family. In this thesis, we adopt the parametric approach and investigate the use of a model family called latent tree models for the task of density estimation. Latent tree models are tree-structured Bayesian networks in which leaf nodes represent observed variables, while internal nodes represent hidden variables. Such models can represent complex relationships among observed variables, and in the meantime, admit efficient inference among them. Consequently, they are a desirable tool for density estimation. While latent tree models are studied for the first time in this thesis for the purpose of density estimation, they have been investigated earlier for clustering and latent structure discovery. Several algorithms for learning latent tree models have been proposed. The state-of-the-art is an algorithm called EAST. EAST determines model structures through principled and systematic search, and determines model parameters using the EM algorithm. It has been shown to be capable of achieving good trade-off between fit to data and model complexity. It is also capable of discovering latent structures behind data. Unfortunately, it has a high computational complexity, which limits its applicability to density estimation problems. In this thesis, we propose two latent tree model learning algorithms specifically for density estimation. The two algorithms have distinct characteristics and are suitable for different applications. The first algorithm is called HCL. HCL assumes a predetermined bound on model complexity and restricts to binary model structures. It first builds a binary tree structure based on mutual information and then runs the EM algorithm once on the resulting structure to determine the parameters. As such, it is efficient and can deal with large applications. The second algorithm is called Pyramid. Pyramid does not assume predetermined bounds on model complexity and does not restrict to binary tree structures. It builds model structures using heuristics based on mutual information and local search. It is slower than HCL. However, it is faster than EAST and is only slightly inferior to EAST in terms of the quality of the resulting models. In this thesis, we also study two applications of the density estimation techniques that we develop. The first application is to approximate probabilistic inference in Bayesian networks. A Bayesian network represents a joint distribution over a set of random variables. It often happens that the network structure is very complex and making inference directly on the network is computational intractable. We propose to approximate the joint distribution using a latent tree model and exploit the latent tree model for faster inference. The idea is to sample data from the Bayesian network, learn a latent tree model from the data offline, and when online, make inference with the latent tree model instead of the original Bayesian network. HCL is used here because the sample size needs to be large to produce accurate approximation and it is possible to predetermine a bound on the online running. Empirical evidence shows that this method can achieve good approximation accuracy at low online computational cost. The second application is classification. A common approach to this task is to formulate it as a density estimation problem: One constructs the class-conditional density for each class and then uses the Bayes rule to make classification. We propose to estimate those class-conditional densities using either EAST or Pyramid. Empiricalevidence shows that this method yields good classification performances. Moreover, the latent tree models built for the class-conditional densities are often meaningful, which is conducive to user confidence. A comparison between EAST and Pyramid reveals that Pyramid is significantly more efficient than EAST, while it results in more or less the same classification performance as the latter.
Read more