- Research Article
66
- 10.1016/j.jspi.2009.03.016
Testing overdispersion in the zero-inflated Poisson model
- Mar 31, 2009
- Journal of Statistical Planning and Inference
- Zhao Yang + 2 more +2
Testing overdispersion in the zero-inflated Poisson model
Zero-inflated count data are characterized by an excessive frequency of zeros that cannot be adequately analyzed by a single distribution, such as Poisson or negative binomial. This problem is pervasive in many practical applications, including document–keyword matrix derived from text corpora, where most keyword frequencies are zero. Conventional statistical approaches, such as the zero-inflated Poisson (ZIP) and zero-inflated negative binomial (ZINB) models, explicitly separate a structural zero component from a count component, but they typically assume independent observations and can be unstable when covariates are high-dimensional and sparse. To address these limitations, this paper proposes a graph-based zero-inflated learning framework that combines simple graph convolution (SGC) with zero-inflated count regression heads such as ZIP and ZINB. We first construct an observation graph by connecting similar samples, and then apply SGC to propagate and smooth features over the graph, producing convolutional representations that incorporate neighborhood information while remaining computationally lightweight. The resulting representations are used as covariates in ZIP and ZINB heads, which preserve probabilistic interpretability through maximum likelihood learning. Our experiments on simulated zero-inflated datasets with controlled zero ratios demonstrate that the proposed ZIP+SGC and ZINB+SGC consistently reduce prediction errors compared with their non-graph baselines, as measured by mean absolute error and root mean squared error. Overall, the proposed approach provides an efficient and interpretable way to integrate graph neural computation with zero-inflated modeling for sparse count prediction problems.
Testing overdispersion in the zero-inflated Poisson model
Testing overdispersion in the zero-inflated Poisson model
Bayesian estimation and case influence diagnostics for the zero-inflated negative binomial regression model
In recent years, there has been considerable interest in regression models based on zero-inflated distributions. These models are commonly encountered in many disciplines, such as medicine, public health, and environmental sciences, among others. The zero-inflated Poisson (ZIP) model has been typically considered for these types of problems. However, the ZIP model can fail if the non-zero counts are overdispersed in relation to the Poisson distribution, hence the zero-inflated negative binomial (ZINB) model may be more appropriate. In this paper, we present a Bayesian approach for fitting the ZINB regression model. This model considers that an observed zero may come from a point mass distribution at zero or from the negative binomial model. The likelihood function is utilized to compute not only some Bayesian model selection measures, but also to develop Bayesian case-deletion influence diagnostics based on q-divergence measures. The approach can be easily implemented using standard Bayesian software, such as WinBUGS. The performance of the proposed method is evaluated with a simulation study. Further, a real data set is analyzed, where we show that ZINB regression models seems to fit the data better than the Poisson counterpart.
Read moreA GEE-type approach to untangle structural and random zeros in predictors.
Count outcomes with excessive zeros are common in behavioral and social studies, and zero-inflated count models such as zero-inflated Poisson (ZIP) and zero-inflated Negative Binomial (ZINB) can be applied when such zero-inflated count data are used as response variable. However, when the zero-inflated count data are used as predictors, ignoring the difference of structural and random zeros can result in biased estimates. In this paper, a generalized estimating equation (GEE)-type mixture model is proposed to jointly model the response of interest and the zero-inflated count predictors. Simulation studies show that the proposed method performs well for practical settings and is more robust for model misspecification than the likelihood-based approach. A case study is also provided for illustration.
Read moreComparison of Zero Inflated Poisson (ZIP) Regression, Zero Inflated Negative Binomial Regression (ZINB) and Binomial Negative Hurdle Regression (HNB) to Model Daily Cigarette Consumption Data for Adult Population in Indonesia
Smoking is a habit that is not good for health. Smoking habits are generally practiced by adults but it is possible for teenagers to do so.The Report of Southeast Asia Tobacco Control Alliance (SEATCA) entitled The Tobacco Control Atlas, ASEAN Region shows that Indonesia is the country with the highest number of smokers in ASEAN, namely 65.19 million people. This figure is equivalent to 34 percent of the total population of Indonesia in 2016. Based on these data, the authors are interested in modeling the daily cigarette consumption data for adults in Indonesia obtained from the 2015 Indonesia Family Life Survey. The variables used include the variable amount of cigarette consumption, education, level of welfare and income per month. The author wants to compare the best model that can be used to model the daily cigarette consumption of adults in Indonesia. The models being compared are Zero Inflated Poisson Regression (ZIP), Zero Inflated Negative Binomial Regression (ZINB) and Binomial Negative Hurdle Regression (HNB). The comparison results of the three models obtained that the best model is the Zero Inflated Negative Binomial (ZINB) Regression model because it has the smallest Akaike's Information Criterion (AIC) value.
Read moreModelling of vertical integration in commercial poultry production of Ghana: A count data model analysis
Modelling of vertical integration in commercial poultry production of Ghana: A count data model analysis
Statistical modelling for falls count data
Statistical modelling for falls count data
Modeling Vehicle-pedestrian Crashes With Excess Zero Along Malaysia Federal Roads
Modeling Vehicle-pedestrian Crashes With Excess Zero Along Malaysia Federal Roads
Evaluating risk factors associated with severe hypoglycaemia in epidemiology studies-what method should we use?
To determine the most appropriate regression models to use when assessing risk factors for severe hypoglycaemia and to investigate the impact of model misspecification and its clinical implications. A total of 1229 children with Type 1 diabetes (mean age 11.7 years sd 4.1), of which 605 (49.2%) were males, were studied. Prospective assessment of severe hypoglycaemia (an event leading to loss of consciousness or seizure) was made over the 9-year period, 1992-2001. Patients were seen every 3 months and episodes of hypoglycaemia along with clinical data were recorded. Over 70% of children never experienced a severe hypoglycaemic event. Data were analysed using the Poisson regression, negative binomial, zero-inflated Poisson (ZIP) and zero-inflated negative binomial (ZINB) models. The over-dispersion and likelihood ratio statistics were calculated and the analytical methods compared. The Poisson regression model did not fit the data well. The negative binomial and the zero inflated Poisson and negative binomial models fitted the data better than Poisson. The commonly used Poisson regression models to analyse hypoglycaemia epidemiology may lead to biased parameter estimates and incorrect determination of risk factors for hypoglycaemia. We recommend the use of the negative binomial or zero inflated models to examine any risk factors associated with severe hypoglycaemia. Careful consideration must be given to the interpretation of hypoglycaemia surveys and their analysis.
Read moreThe Effects of Land Use, Design and Environment on Traffic Fatalities
Developing statistics methods to distinguish significant factors associated with roadways is one of the most feasible accesses to understand the nature of traffic accidents. In this study, zero-inflated negative binomial (ZINB) model was developed to allow for overdispersion and excess zeros, as well as the factors of land use, design and environment to examine the effects. The statistical tests show that ZINB model is preferred to zero-inflated Poisson and negative binomial models due to its ability to describe crash counts associated with severe injuries and fatalities more effectively. The results show that fatalities are positively associated with segment length, surface width, land use variables and rainfall. For example, an increase of one inch rainfall will result in an increase of 0.02% in fatalities. Interestingly, distances to hospitals yield positive impact, which suggests that longer distances lead to higher fatalities, presumably due to time lost in transporting crash victims to hospitals.
Read moreModeling annualized occurrence, frequency, and composition of ingrowth using mixed-effects zero-inflated models and permanent plots in the Acadian Forest Region of North America
Forest tree ingrowth is a highly variable and largely stochastic process. Consequently, predicting occurrence, frequency, and composition of ingrowth is a challenging task but of great importance in long-term forest growth and yield model projections. However, ingrowth data often require different statistical techniques other than traditional Gaussian regression, because these data are often bounded, skewed, and non-normal and commonly contain a large fraction of zeros. This study presents a set of regression models based on discrete Poisson and negative binomial probability distributions for ingrowth data collected from permanent sample plots in the Acadian Forest Region of North America. Models considered here include regular Poisson, zero-inflated Poisson (ZIP), zero-altered Poisson (ZAP; hurdle Poisson), regular negative binomial (NB), zero-inflated negative binomial (ZINB), and zero-altered negative binomial (ZANB; hurdle NB). Plot-level random effects were incorporated into each of these models. The ZINB model with random effects was found to provide the best fit statistics for modeling annualized occurrence and frequency of ingrowth. The key explanatory variables were stand basal area per hectare, percentage of hardwood basal area, number of trees per hectare, a measure of site quality, and the minimum measured diameter at breast height of each plot. A similar model was developed to predict species composition. All models showed logical behavior despite the high variability observed in the original data.
Read moreModeling the Demand for Shared E-Scooter Services
This paper presents the findings on modeling the demand for shared e-scooter services (SES); specifically, spatio-temporal variation of SES demand. A zero-inflated negative binomial (ZINB) model is developed using the count data of trip origins at the dissemination area level from Kelowna, Canada. The motivation for adopting the ZINB model is the presence of excess zeros in the count data. ZINB has two components: the zero-inflated component accounts for excess zeros, and the count component accounts for the over-dispersion characteristics of data resulting from excess zeros. In addition to the ZINB, several other count models including hurdle models are estimated. The goodness-of-fit measures suggest that the ZINB model outperforms other methods. The model results confirm the effects of temporal, weather, transportation infrastructure, land use, and neighborhood characteristics. For example, the count model results reveal that SES demand is more likely to be higher during summer, mid-day on weekends, afternoons of weekdays, and days without rainfall. Furthermore, higher e-scooter index, higher density of cycle tracks, heterogeneous land use, urban centers, lower elevation, and neighborhoods with higher density of hotels and younger population might induce higher demand. The zero component results of the model are consistent with the findings revealed by the count component. The model is validated using a hold-out sample, and the validation results confirm that the prediction performance of the model is reasonably satisfactory. The findings of this study provide important insights into when and where the demand is higher, which will assist in effective policy-making supporting e-scooter use.
Read moreMulti-Task CNN-LSTM Modeling of Zero-Inflated Count and Time-to-Event Outcomes for Causal Inference with Functional Representation of Features
We propose a novel deep learning framework for counterfactual inference on the COMPAS dataset, utilizing a multi-task CNN-LSTM architecture. The model jointly predicts multiple outcome types: (i) count outcomes with zero inflation, modeled using zero-inflated Poisson (ZIP), zero-inflated negative binomial (ZINB), and negative binomial (NB) distributions; (ii) time-to-event outcomes, modeled via the Cox proportional hazards model. To effectively leverage the structure in high-dimensional tabular data, we integrate functional data analysis (FDA) techniques by transforming covariates into smooth functional representations using B-spline basis expansions. Specifically, we construct a pseudo-temporal index over predictor variables and fit basis expansions to each subject’s feature vector, yielding a low-dimensional set of coefficients that preserve smooth variation while reducing noise. This functional representation enables the CNN-LSTM model to capture both local and global temporal patterns in the data, including treatment-covariate interactions. Our approach estimates both population-average and individual-level treatment effects (ATE and CATE) for each outcome and evaluates predictive performance using metrics such as Poisson deviance, root mean squared error (RMSE), and the concordance index (C-index). Statistical inference on treatment effects is supported via bootstrap-based confidence intervals and hypothesis testing. Overall, this comprehensive framework facilitates flexible modeling of heterogeneous treatment effects in structured, high-dimensional data, advancing causal inference methodologies in criminal justice and related domains.
Read more<b>Modeling citrus huanglongbing data using a zero-inflated negative binomial distribution
Zero-inflated data from field experiments can be problematic, as these data require the use of specific statistical models during the analysis process. This study utilized the zero-inflated negative binomial (ZINB) model with the log- and logistic-link functions to describe the incidence of plants with Huanglongbing (HLB, caused by Candidatus liberibacter spp.) in commercial citrus orchards in the Northwestern Parana State, Brazil. Each orchard was evaluated at different times. The ZINB model with random effects in both link functions provided the best fit, as the inclusion of these effects accounted for variations between orchards and the numbers of diseased plants. The results of this model show that older plants exhibit a lower probability of acquiring HLB. The application of insecticides on a calendar basis or during new foliage flushes resulted in a three times larger probability of developing HLB compared with applying insecticides only when the vector was detected.
Read moreZERO-INFLATED POISSON REGRESSION MODELS: APPLICATIONS IN THE SCIENCES AND SOCIAL SCIENCES
This paper makes a theoretical contribution by presenting a detailed derivation of a zero-inflated Poisson (ZIP) model, and then deriving the parameters of the ZIP model using a fishing data set. This model has several practical applications, and is largely performed to model count data that have an excess number of zero counts. In the scope of the paper, we introduce the complete formulae, the likelihood and log-likelihood functions and the estimating equation of the ZIP model. We then investigate the theory of large sample properties of this model under some regularity conditions. A simulation study and a fishing data set are studied for the ZIP model. The results in the actual application in this work are meaningful, useful and crucial in reality. The results also provide reliable evidence for obtaining the largest number of fish while fishing. This is the contribution of this research in terms of applications. Finally, the important applications of this model in practice, some conclusions, and future work is also presented for consideration.
Read moreSpatial Scan Statistics for Models with Excess Zeros and Overdispersion
ObjectiveTo propose a more realistic model for disease cluster detection, through a modification of the spatial scan statistic to account simultaneously for inflated zeros and overdispersion.IntroductionSpatial Scan Statistics [1] usually assume Poisson or Binomial distributed data, which is not adequate in many disease surveillance scenarios. For example, small areas distant from hospitals may exhibit a smaller number of cases than expected in those simple models. Also, underreporting may occur in underdeveloped regions, due to inefficient data collection or the difficulty to access remote sites. Those factors generate excess zero case counts or overdispersion, inducing a violation of the statistical model and also increasing the type I error (false alarms). Overdispersion occurs when data variance is greater than the predicted by the used model. To accommodate it, an extra parameter must be included; in the Poisson model, one makes the variance equal to the mean.MethodsTools like the Generalized Poisson (GP) and the Double Poisson [2] may be a better option for this kind of problem, modeling separately the mean and variance, which could be easily adjusted by covariates. When excess zeros occur, the Zero Inflated Poisson (ZIP) model is used, although ZIP’s estimated parameters may be severely biased if nonzero counts are too dispersed, compared to the Poisson distribution. In this case the Inflated Zero models for the Generalized Poisson (ZIGP), Double Poisson (ZIDP) and Negative Binomial (ZINB) could be good alternatives to the joint modeling of excess zeros and overdispersion. By one hand, Zero Inflated Poisson (ZIP) models were proposed using the spatial scan statistic to deal with the excess zeros [3]. By the other hand, another spatial scan statistic was based on a Poisson-Gamma mixture model for overdispersion [4]. In this work we present a model which includes inflated zeros and overdispersion simultaneously, based on the ZIDP model. Let the parameter p indicate the zero inflation. As the the remaining parameters of the observed cases map and the parameter p are not independent, the likelihood maximization process is not straightforward; it becomes even more complicated when we include covariates in the analysis. To solve this problem we introduce a vector of latent variables in order to factorize the likelihood, and obtain a facilitator for the maximization process using the E-M (Expectation-Maximization) algorithm. We derive the formulas to maximize iteratively the likelihood, and implement a computer program using the E-M algorithm to estimate the parameters under null and alternative hypothesis. The p-value is obtained via the Fast Double Bootstrap Test [5].ResultsNumerical simulations are conducted to assess the effectiveness of the method. We present results for Hanseniasis surveillance in the Brazilian Amazon in 2010 using this technique. We obtain the most likely spatial clusters for the Poisson, ZIP, Poisson-Gamma mixture and ZIDP models and compare the results.ConclusionsThe Zero Inflated Double Poisson Spatial Scan Statistic for disease cluster detection incorporates the flexibility of previous models, accounting for inflated zeros and overdispersion simultaneously.The Hanseniasis study case map, due to excess of zero cases counts in many municipalities of the Brazilian Amazon and the presence of overdispersion, was a good benchmark to test the ZIDP model. The results obtained are easier to understand compared to each of the previous spatial scan statistic models, the Zero Inflated Poisson (ZIP) model and the Poisson-Gamma mixture model for overdispersion, taken separetely. The E-M algorithm and the Fast Double Bootstrap test are computationally efficient for this type of problem.
Read more