- Research Article
97
- 10.1016/j.eswa.2013.11.025
Density weighted support vector data description
- Dec 04, 2013
- Expert Systems with Applications
- Myungraee Cha + 2 more +2
Density weighted support vector data description
Information entropy based sample reduction for support vector data description
Density weighted support vector data description
Density weighted support vector data description
Towards support vector data description based on heuristic sample condensed rule
Support vector data description (SVDD) is a well-known kernel-based one-class classification method that exhibits intrinsic regularization ability and robustness versus low numbers of high-dimensional samples. However, the efficiency of SVDD is limited by the cubic time complexity. To solve this problem, this paper first investigates the effect of selecting a reduced subset as the training set of SVDD, while guaranteeing the classification quality. To this end, a new heuristic sample condensed rule, termed HSC, is proposed to accurately identify those potential support vectors that characterize the classification boundary. HSC can consider both the spatial distribution and local density features of training samples, and focus on selecting samples very close to the decision boundary. When dealing with the local density computation, we introduce the idea of K nearest neighbors (KNN) to examine the density of samples in the neighbors of the object to be classified. Finally, a condensed but informative subset obtained by HSC will be applied to train SVDD breezily. The experimental results show that HSC-based SVDD sensibly improves over conventional SVDD, in terms of the size of the training set while guaranteeing a comparable classification quality. In addition, it is competitive over other improved SVDD classifiers in terms of training and testing time.
Read moreDiagnostic Pattern Recognition on Gene-Expression Profile Data by Using One-Class Classification
In this paper, we perform diagnostic pattern recognition on a gene-expression profile data set by using one-class classification. Unlike conventional multiclass classifiers, the one-class (OC) classifier is built on one class only. For optimal performance, it accepts samples coming from the class used for training and rejects all samples from other classes. We evaluate six OC classifiers: the Gaussian model, Parzen windows, support vector data description (with two types of kernels: inner product and Gaussian), nearest neighbor data description, K-means, and PCA on three gene-expression profile data sets, those being an SRBCT data set, a Colon data set, and a Leukemia data set. Providing there is a good splitting of training and test samples and feature selection, most OC classifiers can produce high quality results. Parzen windows and support vector data description are "over-strict" in most cases, while nearest neighbor data description is "over-loose". Other classifiers are intermediate between these two extremes. The main difficulty for the OC classifier is it is difficult to obtain an optimum decision threshold if there are a limited number of training samples.
Read moreA comparative investigation of data-driven approaches based on one-class classifiers for condition monitoring of marine machinery system
A comparative investigation of data-driven approaches based on one-class classifiers for condition monitoring of marine machinery system
Read moreBoundary‐based Fuzzy‐SVDD for one‐class classification
Support Vector Data Description (SVDD) is an extremely hot topic issue in One-Class Classification (OCC), which has displayed outstanding performance in dealing with many novelty detection problems. However, SVDD just takes the data description by the kernel-based distance among each instance into consideration rather than considering the distribution of the data. Therefore, Fuzzy Support Vector Data Description (Fuzzy-SVDD) has been developed to distribute a fuzzy membership to each input sample so that different samples cause different contributions to classification boundary. The majority of the methods in Fuzzy-SVDD are based on the sample density, but there are remaining two problems. These density-based Fuzzy-SVDD methods would decrease the contribution of support vectors (SVs) in low densities. What is more, these methods cannot get a precise density when there are few target samples. These two problems would lead to a poor classification boundary. To overcome these drawbacks, a novel method called Boundary-based Fuzzy-SVDD (BF-SVDD) is proposed in this paper. BF-SVDD uses a new definition called local–global center distance to search for the samples near the boundary. Then, it enhances fuzzy memberships of these samples because they carry more significant information for the decision boundary than other data. The contribution of this paper can be summarized into three main points. First a novel concept called local–global center distances is proposed to find the SVs better. Second, fuzzy memberships with local–global center distance make SVs more informative to create the decision boundary. Furthermore, the experiments based on University of California, Irvine and Knowledge Extraction based on Evolutionary Learning also show that the proposed method has excellent performances. Even for the minority class in imbalance data sets, the proposed method can also have a good classification.
Read moreClass-Incremental Learning Based on Feature Extraction of CNN With Optimized Softmax and One-Class Classifiers
With the development of deep convolutional neural networks in recent years, the network structure has become more and more complicated and varied, and there are very good results in pattern recognition, image classification, scene classification, and target tracking. This end-to-end learning model relies on the initial large dataset. However, many data are gradually obtained in practical situations, which contradict the deep learning of one-time batch learning. There is an urgent need for an incremental learning approach that can continuously learn new knowledge from new data while retaining what has already been learned. This paper proposes an incremental learning algorithm based on convolutional neural network and support vector data description. CNN and AM-Softmax loss function are used to represent and continuously learn image features. Support vector data description is used to construct multiple hyperspheres for new and old classes of images. Class-incremental learning is achieved by the increment of hyperspheres. The experimental results show that the incremental learning method proposed in this paper can effectively extract the latent features of the image and adapt it to the learning situation of the class-increment. The recognition accuracy is close to batch learning.
Read moreMonitoring of solid-state fermentation of wheat straw in a pilot scale using FT-NIR spectroscopy and support vector data description
Monitoring of solid-state fermentation of wheat straw in a pilot scale using FT-NIR spectroscopy and support vector data description
Read moreRisk Assessment Model Based on SVDD and Fuzzy Regression Method
The paper aims to solve the problem of insufficient high risk data in risk assessment of R&D projects. A one-class classification method called support vector data description (SVDD) is studied, and an intelligent risk assessment model based on SVDD with fuzzy regression information is also proposed. The model comes into being a new approach. Applying this approach, firstly verify the conversional risk evaluation indexes by fuzzy regression technique to develop a sensitive index system. Secondly the study uses the historical risk data referring to these indexes to train the SVDD one-class classifier. Unlike previously proposed intelligent methods of risk assessment, with this model the risk level can be distinguished only by training of low risk data. The results of its application on an example show that the method is feasible for risk assessment with the fuzzy high risk data.
Read moreEdge-pixels-based support vector data description for specific land-cover distribution mapping
An edge-pixels-based support vector data description (EPSVDD) method has been developed for improving one-class classification accuracy. The proposed method was validated in two experiments: a simulated experiment and an actual experiment. In the simulated experiment, a ring segmentation search method was performed to segment the wheat spectral feature for deriving training samples of different spectral responses. As the training data moved from the center to the edge of wheat distribution, the hypersphere expanded and the overall accuracy (OA) simultaneously increased, highlighting the potential advantage of edge pixels in SVDD classification. In the actual experiment, edge training samples were manually acquired from geographical parcel boundary and minimum noise fraction (MNF) scatterplots for both wheat and bare-land classes. For the wheat class, EPSVDD yielded an improved classification with an OA of 92.71% and a producer’s accuracy of 95.81%, which were higher than those of conventional SVDD method using typical training samples. Similarly, for the bare-land class, the OA of the EPSVDD was 92.53%, which was also significantly higher than traditional SVDD method. Then, SVDD classifications were carried out and repeated 10 times using different training set sizes. Mean OAs were almost higher than 0.9 with variance less than 0.03 using edge training samples, while highest OAs for wheat and bare land classes were 0.74 and 0.81, respectively, using random sampling method. The EPSVDD can effectively select the informative training sample for SVDD classifier to improve the accuracy of one-class classification.
Read moreDeep learning with support vector data description
Deep learning with support vector data description
Active Learning of SVDD Hyperparameter Values
Support Vector Data Description (SVDD) is a popular one-class classifier, and well-suited for outlier detection. However, the effectiveness of SVDD depends on selecting good hyperparameter values – a difficult problem that has received significant attention in the literature. Since SVDD is an unsupervised classifier, tuning of hyperparameter values is difficult. This has motivated several methods to estimate hyperparameter values based on data characteristics. But existing methods are purely heuristic, and the conditions under which they work well are largely unclear. This has created a situation where instead of selecting hyperparameter values, one has to choose among several, equally plausible heuristics.In this article, we make some strides towards a principled approach to estimate SVDD hyperparameter values. We propose LAMA (Local Active Min-Max Alignment), the first method to select SVDD hyperparameter values by active learning. The core idea is based on kernel alignment, which we adapt to active learning with small sample sizes. LAMA provides estimates for both of the SVDD hyperparameters. These estimates are evidence-based, i.e., rely on actual class labels, and come with a quality score. This eliminates the need for manual validation, an issue with current heuristics. LAMA outperforms state-of-theart competitors in extensive experiments on real-world data. In several cases, LAMA even yields results close to the empirical upper bound.
Read moreAn Adaptive Radial Basis Function Kernel for Support Vector Data Description
For one-class classification or novelty detection, the metric of the feature space is essential for a good performance. Typically, it is assumed that the metric of the feature space is relatively isotropic, or flat, indicating that a distance of 1 can be interpreted in a similar way for every location and direction in the feature space. When this is not the case, thresholds on distances that are fitted in one part of the feature space will be suboptimal for other parts. To avoid this, the idea of this paper is to modify the width parameter in the Radial Basis Function (RBF) kernel for the Support Vector Data Description (SVDD) classifier. Although there have been numerous approaches to learn the metric in a feature space for (supervised) classification problems, for one-class classification this is harder, because the metric cannot be optimized to improve a classification performance. Instead, here we propose to consider the local pairwise distances in the training set. The results obtained on both artificial and real datasets demonstrate the ability of the modified RBF kernel to identify local scales in the input data, extracting its general structure and improving the final classification performance for novelty detection problems.
Read moreMonitoring the mean with least-squares support vector data description
Abstract: Multivariate control charts are essential tools in multivariate statistical process control (MSPC). “Shewhart-type” charts are control charts using rational subgroupings which are effective in the detection of large shifts. Recently, the one-class classification problem has attracted a lot of interest. Three methods are typically used to solve this type of classification problem. These methods include the k−center method, the nearest neighbor method, one-class support vector machine (OCSVM), and the support vector data description (SVDD). In industrial applications, like statistical process control (SPC), practitioners successfully used SVDD to detect anomalies or outliers in the process. In this paper, we reformulate the standard support vector data description and derive a least squares version of the method. This least-squares support vector data description (LS-SVDD) is used to design a control chart for monitoring the mean vector of processes. We compare the performance of the LS-SVDD chart with the SVDD and T2 chart using out-of-control Average Run Length (ARL) as the performance metric. The experimental results indicate that the proposed control chart has very good performance.
Read moreReducing the Impact of Outliers on the One-Class Classification Decision Rule
A modified version of one-class classification criterion reducing the impact of outliers on the one-class classification decision rule is proposed based on support vector data description (SVDD) by D. Tax. The optimization method utilizes the substitution of nondifferentiable objective function by the smooth one. A comparative experimental study of existing one-class methods shows the superiority of the proposed criterion in anomaly detection.
Read moreVerification-Based Design of a Robust EMG Wake Word.
Surface electromyography (sEMG) signals are now commonly used in continuous myoelectric control of prostheses. More recently, researchers have considered EMG-based gesture recognition systems for human computer interaction research. These systems instead focus on recognizing discrete gestures (like a finger snap). The majority of works, however, have focused on improving multi-class performance, with little consideration for false activations from "other" classes. Consequently, they lack the robustness needed for real-world applications which generally require a single motion class such as a mouse click or a wake word. Furthermore, many works have borrowed the windowed classification schemes from continuous control, and thus fail to leverage the temporal structure of the gesture. In this paper, we propose a verification-based approach to creating a robust EMG wake word using one-class classifiers (Support Vector Data Description, One Class-Support Vector Machine, Dynamic Time Warping (DTW) & Hidden Markov Models). The area under the ROC curve (AUC) is used as a feature optimization objective as it provides a better representation of the verification performance. Equal error rate (EER) and AUC are then used as evaluation metrics. The results are computed using both window-based and temporal classifiers on a dataset consisting of five different gestures, with a best EER of 0.04 and AUC of 0.98, recorded using a DTW scheme. These results demonstrate a design framework that may benefit the development of more robust solutions for EMG-based wake words or input commands for a variety of interactive applications.
Read more