- Research Article
3
- 10.1016/s0024-3795(01)00425-6
Methods of density estimation on the Grassmann manifold
- Sep 04, 2002
- Linear Algebra and its Applications
- Yasuko Chikuse
Methods of density estimation on the Grassmann manifold
We propose a method for nonparametric density estimation that exhibits robustness to contamination of the training sample. This method achieves robustness by combining a traditional kernel density estimator (KDE) with ideas from classical M-estimation. We interpret the KDE based on a positive semi-definite kernel as a sample mean in the associated reproducing kernel Hilbert space. Since the sample mean is sensitive to outliers, we estimate it robustly via M-estimation, yielding a robust kernel density estimator (RKDE). An RKDE can be computed efficiently via a kernelized iteratively re-weighted least squares (IRWLS) algorithm. Necessary and sufficient conditions are given for kernelized IRWLS to converge to the global minimizer of the M-estimator objective function. The robustness of the RKDE is demonstrated with a representer theorem, the influence function, and experimental results for density estimation and anomaly detection.
Methods of density estimation on the Grassmann manifold
Methods of density estimation on the Grassmann manifold
Spline local basis methods for nonparametric density estimation
This work reviews the literature on spline local basis methods for non-parametric density estimation. Particular attention is paid to B-spline density estimators which have experienced recent advances in both theory and methodology. These estimators occupy a very interesting space in statistics, which lies aptly at the cross-section of numerous statistical frameworks. New insights, experiments, and analyses are presented to cast the various estimation concepts in a unified context, while parallels and contrasts are drawn to the more familiar contexts of kernel density estimation. Unlike kernel density estimation, the study of local basis estimation is not yet fully mature, and this work also aims to highlight the gaps in existing literature which merit further investigation.
Read moreA note on nonparametric density deconvolution by weighted kernel estimators
Recently Hazelton and Turlach (2009) proposed a weighted kernel density estimatorfor the deconvolution problem. In the case of Gaussian kernels and measurement er-ror, they argued that the weighted kernel density estimator is a competitive estimatorover the classical deconvolution kernel estimator. In this paper we consider weightedkernel density estimators when sample observations are contaminated by double expo-nentially distributed errors. The performance of the weighted kernel density estimatorsis compared over the classical deconvolution kernel estimator and the kernel density es-timator based on the support vector regression method by means of a simulation study.The weighted density estimator with the Gaussian kernel shows numerical instabilityin practical implementation of optimization function. However the weighted densityestimates with the double exponential kernel has very similar patterns to the classicalkernel density estimates in the simulations, but the shape is less satisfactory than theclassical kernel density estimator with the Gaussian kernel.Keywords: Deconvolution, kernel density estimator, support vector regression, weightedkernel density estimator.
Read moreDm-KDE: dynamical kernel density estimation by sequences of KDE estimators with fixed number of components over data streams
In many data stream mining applications, traditional density estimation methods such as kernel density estimation, reduced set density estimation can not be applied to the density estimation of data streams because of their high computational burden, processing time and intensive memory allocation requirement. In order to reduce the time and space complexity, a novel density estimation method Dm-KDE over data streams based on the proposed algorithm m-KDE which can be used to design a KDE estimator with the fixed number of kernel components for a dataset is proposed. In this method, Dm-KDE sequence entries are created by algorithm m-KDE instead of all kernels obtained from other density estimation methods. In order to further reduce the storage space, Dm-KDE sequence entries can be merged by calculating their KL divergences. Finally, the probability density functions over arbitrary time or entire time can be estimated through the obtained estimation model. In contrast to the state-of-the-art algorithm SOMKE, the distinctive advantage of the proposed algorithm Dm-KDE exists in that it can achieve the same accuracy with much less fixed number of kernel components such that it is suitable for the scenarios where higher on-line computation about the kernel density estimation over data streams is required.We compare Dm-KDE with SOMKE and M-kernel in terms of density estimation accuracy and running time for various stationary datasets. We also apply Dm-KDE to evolving data streams. Experimental results illustrate the effectiveness of the proposed method.
Read moreQuasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities
Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 106 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.
Read moreNonparametric estimation of Fisher information from real data.
The Fisher information matrix (FIM) is a widely used measure for applications including statistical inference, information geometry, experiment design, and the study of criticality in biological systems. The FIM is defined for a parametric family of probability distributions and its estimation from data follows one of two paths: either the distribution is assumed to be known and the parameters are estimated from the data or the parameters are known and the distribution is estimated from the data. We consider the latter case which is applicable, for example, to experiments where the parameters are controlled by the experimenter and a complicated relation exists between the input parameters and the resulting distribution of the data. Since we assume that the distribution is unknown, we use a nonparametric density estimation on the data and then compute the FIM directly from that estimate using a finite-difference approximation to estimate the derivatives in its definition. The accuracy of the estimate depends on both the method of nonparametric estimation and the difference Δθ between the densities used in the finite-difference formula. We develop an approach for choosing the optimal parameter difference Δθ based on large deviations theory and compare two nonparametric density estimation methods, the Gaussian kernel density estimator and a novel density estimation using field theory method. We also compare these two methods to a recently published approach that circumvents the need for density estimation by estimating a nonparametric f divergence and using it to approximate the FIM. We use the Fisher information of the normal distribution to validate our method and as a more involved example we compute the temperature component of the FIM in the two-dimensional Ising model and show that it obeys the expected relation to the heat capacity and therefore peaks at the phase transition at the correct critical temperature.
Read moreSmoothing level selection for density estimators based on the moments
This paper introduces an approach to select the bandwidth or smoothing parameter in multiresolution (MR) density estimation and nonparametric density estimation. It is based on the evolution of the second, third and fourth central moments and the shape of the estimated densities for different bandwidths and resolution levels. The proposed method has been applied to density estimation by means of multiresolution densities as well as kernel density estimation (MRDE and KDE respectively). The results of the simulations and the empirical application demonstrate that the level of resolution resulting from the moments method performs better with multimodal densities than the Bayesian Information Criterion (BIC) for multiresolution densities estimation and the plug-in for kernel densities estimation.
Read moreComparative Evaluation of Nonparametric Density Estimators for Gaussian Mixture Models with Clustering Support
The article investigates the accuracy of nonparametric univariate density estimation methods applied to various Gaussian mixture models. A comprehensive comparative analysis is performed for four popular estimation approaches: adaptive kernel density estimation, projection pursuit, log-spline estimation, and wavelet-based estimation. The study is extended with modified versions of these methods, where the sample is first clustered using the EM algorithm based on Gaussian mixture components prior to density estimation. Estimation accuracy is quantitatively evaluated using MAE and MAPE criteria, with simulation experiments conducted over 100,000 replications for various sample sizes. The results show that estimation accuracy strongly depends on the density structure, sample size, and degree of component overlap. Clustering before density estimation significantly improves accuracy for multimodal and asymmetric densities. Although no formal statistical tests are conducted, the performance improvement is validated through non-overlapping confidence intervals obtained from 100,000 simulation replications. In addition, several decision-making systems are compared for automatically selecting the most appropriate estimation method based on the sample’s statistical features. Among the tested systems, kernel discriminant analysis yielded the lowest error rates, while neural networks and hybrid methods showed competitive but more variable performance depending on the evaluation criterion. The findings highlight the importance of using structurally adaptive estimators and automation of method selection in nonparametric statistics. The article concludes with recommendations for method selection based on sample characteristics and outlines future research directions, including extensions to multivariate settings and real-time decision-making systems.
Read moreA Support Vector Method for the Deconvolution Problem
This paper considers the problem of nonparametric deconvolution density estimation when sample observa-tions are contaminated by double exponentially distributed errors. Three different deconvolution density estima-tors are introduced: a weighted kernel density estimator, a kernel density estimator based on the support vector regression method in a RKHS, and a classical kernel density estimator. The performance of these deconvolution density estimators is compared by means of a simulation study.
Read moreFully Data-driven Normalized and Exponentiated Kernel Density Estimator with Hyvärinen Score
Recently, Jewson and Rossell (2022) proposed a new approach for kernel density estimation using an exponentiated form of kernel density estimators. The density estimator contained two hyperparameters that flexibly controls the smoothness of the resulting density. We tune them in a data-driven manner by minimizing an objective function based on the Hyvärinen score to avoid the optimization involving the intractable normalizing constant caused by the exponentiation. We show the asymptotic properties of the proposed estimator and emphasize the importance of including the two hyperparameters for flexible density estimation. Our simulation studies and application to income data show that the proposed density estimator is promising when the underlying density is multi-modal or when observations contain outliers.
Read moreFast Kernel Density Estimation with Density Matrices and Random Fourier Features
Kernel density estimation (KDE) is one of the most widely used nonparametric density estimation methods. The fact that it is a memory-based method, i.e., it uses the entire training data set for prediction, makes it unsuitable for most current big data applications. Several strategies, such as tree-based or hashing-based estimators, have been proposed to improve the efficiency of the kernel density estimation method. The novel density kernel density estimation method (DMKDE) uses density matrices, a quantum mechanical formalism, and random Fourier features, an explicit kernel approximation, to produce density estimates. This method has its roots in the KDE and can be considered as an approximation method, without its memory-based restriction. In this paper, we systematically evaluate the novel DMKDE algorithm and compare it with other state-of-the-art fast procedures for approximating the kernel density estimation method on different synthetic data sets. Our experimental results show that DMKDE is on par with its competitors for computing density estimates and advantages are shown when performed on high-dimensional data. We have made all the code available as an open source software repository.
Read moreSmoothDE: a smooth density estimator with good performance
Probability density estimation is the problem of inferring an underlying probability from a sampling of points. This study introduces smoothDE, an algorithm that uses Bayesian Field Theory to optimize non-parametric density estimation. smoothDE deterministically finds an optimal density function based on the probability of observed data, subject to a smoothing constraint and its associated prior probability. smoothDE's predicted densities have almost universally lower Kullback-Leibler divergences from simulated Gaussian Mixtures densities when compared to similar Bayesian Field Theory methods and Kernel Density Estimators. smoothDE was even able to outperform a specialized Bayesian Gaussian Mixture density estimator at lower samplings. smoothDE's ability to quickly fit arbitrary densities allowed it to be used as a preprocessing step for classification algorithm, in certain cases boosting classifier performance.
Read moreMonitoring Non-normal Data with Principal Component Analysis and Adaptive Density Estimation
The issue of monitoring non-normally distributed data with principal component analysis (PCA) is addressed through the application of density estimation for evaluating the quality of the principal component scores. Although kernel density estimation has been previously cited as a method for monitoring such data, mixture models are proposed here in order to reduce model complexity and computational effort. Furthermore, several adaptation strategies for the density estimators are developed and suggestions are provided on their use. A rapid thermal anneal case study demonstrates how the estimators outperform the traditional Hotelling's T <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> statistic due to the presence of a first wafer effect.
Read moreKernel density smoothing of composite spatial data on administrative area level
Composite spatial data on administrative area level are often presented by maps. The aim is to detect regional differences in the concentration of subpopulations, like elderly persons, ethnic minorities, low-educated persons, voters of a political party or persons with a certain disease. Thematic collections of such maps are presented in different atlases. The standard presentation is by Choropleth maps where each administrative unit is represented by a single value. These maps can be criticized under three aspects: the implicit assumption of a uniform distribution within the area, the instability of the resulting map with respect to a change of the reference area and the discontinuities of the maps at the borderlines of the reference areas which inhibit the detection of regional clusters.In order to address these problems we use a density approach in the construction of maps. This approach does not enforce a local uniform distribution. It does not depend on a specific choice of area reference system and there are no discontinuities in the displayed maps. A standard estimation procedure of densities are Kernel density estimates. However, these estimates need the geo-coordinates of the single units which are not at disposal as we have only access to the aggregates of some area system. To overcome this hurdle, we use a statistical simulation concept. This can be interpreted as a Simulated Expectation Maximisation (SEM) algorithm of Celeux et al (1996). We simulate observations from the current density estimates which are consistent with the aggregation information (S-step). Then we apply the Kernel density estimator to the simulated sample which gives the next density estimate (E-Step).This concept has been first applied for grid data with rectangular areas, see Groß et al (2017), for the display of ethnic minorities. In a second application we demonstrated the use of this approach for the so-called “change of support” (Bradley et al 2016) problem. Here Groß et al (2020) used the SEM algorithm to recalculate case numbers between non-hierarchical administrative area systems. Recently Rendtel et al (2021) applied the SEM algorithm to display spatial-temporal clusters of Corona infections in Germany.Here we present three modifications of the basic SEM algorithm: 1) We introduce a boundary correction which removes the underestimation of kernel density estimates at the borders of the population area. 2) We recognize unsettled areas, like lakes, parks and industrial areas, in the computation of the kernel density. 3) We adapt the SEM algorithm for the computation of local percentages which are important especially in voting analysis.We evaluate our approach against several standard maps by means of the local voting register with known addresses. In the empirical part we apply our approach for the display of voting results for the 2016 election of the Berlin parliament. We contrast our results against Choropleth maps and show new possibilities for reporting spatial voting results.
Read moreBayesian nonparametric estimation of bandwidth using mixtures of kernel estimators for length-biased data
ABSTRACTKernel density estimation has been applied in many computational subjects. In this paper, we propose a density estimation procedure from a Bayesian nonparametric perspective using Dirichlet process prior for the length-biased data under an unknown kernel function. In this situation, the kernel within the Dirichlet process mixture model will be approximated by the kernel density estimator. We present a Bayesian nonparametric method for finding the bandwidth parameter in the kernel density estimation using a Markov chain Monte Carlo approach. Then, this approach is used to the simulated and real data set. Finally, we compare the proposed bandwidth estimation with the other estimations like cross-validation and Bayes based on the mean integrated squared error criterion.
Read more