- Research Article
97
- 10.1016/j.eswa.2013.11.025
Density weighted support vector data description
- Dec 04, 2013
- Expert Systems with Applications
- Myungraee Cha + 2 more +2
Density weighted support vector data description
In this paper, we perform diagnostic pattern recognition on a gene-expression profile data set by using one-class classification. Unlike conventional multiclass classifiers, the one-class (OC) classifier is built on one class only. For optimal performance, it accepts samples coming from the class used for training and rejects all samples from other classes. We evaluate six OC classifiers: the Gaussian model, Parzen windows, support vector data description (with two types of kernels: inner product and Gaussian), nearest neighbor data description, K-means, and PCA on three gene-expression profile data sets, those being an SRBCT data set, a Colon data set, and a Leukemia data set. Providing there is a good splitting of training and test samples and feature selection, most OC classifiers can produce high quality results. Parzen windows and support vector data description are "over-strict" in most cases, while nearest neighbor data description is "over-loose". Other classifiers are intermediate between these two extremes. The main difficulty for the OC classifier is it is difficult to obtain an optimum decision threshold if there are a limited number of training samples.
Density weighted support vector data description
Density weighted support vector data description
Boundary‐based Fuzzy‐SVDD for one‐class classification
Support Vector Data Description (SVDD) is an extremely hot topic issue in One-Class Classification (OCC), which has displayed outstanding performance in dealing with many novelty detection problems. However, SVDD just takes the data description by the kernel-based distance among each instance into consideration rather than considering the distribution of the data. Therefore, Fuzzy Support Vector Data Description (Fuzzy-SVDD) has been developed to distribute a fuzzy membership to each input sample so that different samples cause different contributions to classification boundary. The majority of the methods in Fuzzy-SVDD are based on the sample density, but there are remaining two problems. These density-based Fuzzy-SVDD methods would decrease the contribution of support vectors (SVs) in low densities. What is more, these methods cannot get a precise density when there are few target samples. These two problems would lead to a poor classification boundary. To overcome these drawbacks, a novel method called Boundary-based Fuzzy-SVDD (BF-SVDD) is proposed in this paper. BF-SVDD uses a new definition called local–global center distance to search for the samples near the boundary. Then, it enhances fuzzy memberships of these samples because they carry more significant information for the decision boundary than other data. The contribution of this paper can be summarized into three main points. First a novel concept called local–global center distances is proposed to find the SVs better. Second, fuzzy memberships with local–global center distance make SVs more informative to create the decision boundary. Furthermore, the experiments based on University of California, Irvine and Knowledge Extraction based on Evolutionary Learning also show that the proposed method has excellent performances. Even for the minority class in imbalance data sets, the proposed method can also have a good classification.
Read moreA comparative investigation of data-driven approaches based on one-class classifiers for condition monitoring of marine machinery system
A comparative investigation of data-driven approaches based on one-class classifiers for condition monitoring of marine machinery system
Read moreInformation entropy based sample reduction for support vector data description
Information entropy based sample reduction for support vector data description
Edge-pixels-based support vector data description for specific land-cover distribution mapping
An edge-pixels-based support vector data description (EPSVDD) method has been developed for improving one-class classification accuracy. The proposed method was validated in two experiments: a simulated experiment and an actual experiment. In the simulated experiment, a ring segmentation search method was performed to segment the wheat spectral feature for deriving training samples of different spectral responses. As the training data moved from the center to the edge of wheat distribution, the hypersphere expanded and the overall accuracy (OA) simultaneously increased, highlighting the potential advantage of edge pixels in SVDD classification. In the actual experiment, edge training samples were manually acquired from geographical parcel boundary and minimum noise fraction (MNF) scatterplots for both wheat and bare-land classes. For the wheat class, EPSVDD yielded an improved classification with an OA of 92.71% and a producer’s accuracy of 95.81%, which were higher than those of conventional SVDD method using typical training samples. Similarly, for the bare-land class, the OA of the EPSVDD was 92.53%, which was also significantly higher than traditional SVDD method. Then, SVDD classifications were carried out and repeated 10 times using different training set sizes. Mean OAs were almost higher than 0.9 with variance less than 0.03 using edge training samples, while highest OAs for wheat and bare land classes were 0.74 and 0.81, respectively, using random sampling method. The EPSVDD can effectively select the informative training sample for SVDD classifier to improve the accuracy of one-class classification.
Read moreBinary classification based on SVDD projection and nearest neighbors
The SVDD (support vector data description) is one of the most well-known one-class support vector learning methods, in which one tries the strategy of utilizing balls defined on the feature space in order to distinguish a set of normal data from all other possible abnormal objects. The usual strategy of the SVDD depends on the process of finding the region for the normal-class training data with somewhat neglecting the distribution of the abnormal data, thus it may not work well if applied to the binary classification problems in which two classes have similar number of data. In this paper, we consider the problem of performing binary classification based on the SVDD techniques, and in order to overcome the possible drawback of the usual SVDD strategy focusing on the normal-class data only, we propose a new SVDD algorithm which is based on the use of two different SVDD balls for the positive and negative classes along with the SVDD projection and nearest neighbor rule. To investigate how the proposed method works, we compared the performance of the proposed method with SVC (support vector classifier) and conventional SVDD using several real datasets.
Read moreClass-Incremental Learning Based on Feature Extraction of CNN With Optimized Softmax and One-Class Classifiers
With the development of deep convolutional neural networks in recent years, the network structure has become more and more complicated and varied, and there are very good results in pattern recognition, image classification, scene classification, and target tracking. This end-to-end learning model relies on the initial large dataset. However, many data are gradually obtained in practical situations, which contradict the deep learning of one-time batch learning. There is an urgent need for an incremental learning approach that can continuously learn new knowledge from new data while retaining what has already been learned. This paper proposes an incremental learning algorithm based on convolutional neural network and support vector data description. CNN and AM-Softmax loss function are used to represent and continuously learn image features. Support vector data description is used to construct multiple hyperspheres for new and old classes of images. Class-incremental learning is achieved by the increment of hyperspheres. The experimental results show that the incremental learning method proposed in this paper can effectively extract the latent features of the image and adapt it to the learning situation of the class-increment. The recognition accuracy is close to batch learning.
Read moreTowards support vector data description based on heuristic sample condensed rule
Support vector data description (SVDD) is a well-known kernel-based one-class classification method that exhibits intrinsic regularization ability and robustness versus low numbers of high-dimensional samples. However, the efficiency of SVDD is limited by the cubic time complexity. To solve this problem, this paper first investigates the effect of selecting a reduced subset as the training set of SVDD, while guaranteeing the classification quality. To this end, a new heuristic sample condensed rule, termed HSC, is proposed to accurately identify those potential support vectors that characterize the classification boundary. HSC can consider both the spatial distribution and local density features of training samples, and focus on selecting samples very close to the decision boundary. When dealing with the local density computation, we introduce the idea of K nearest neighbors (KNN) to examine the density of samples in the neighbors of the object to be classified. Finally, a condensed but informative subset obtained by HSC will be applied to train SVDD breezily. The experimental results show that HSC-based SVDD sensibly improves over conventional SVDD, in terms of the size of the training set while guaranteeing a comparable classification quality. In addition, it is competitive over other improved SVDD classifiers in terms of training and testing time.
Read moreA dynamic ensemble outlier detection model based on an adaptive k-nearest neighbor rule
A dynamic ensemble outlier detection model based on an adaptive k-nearest neighbor rule
Reducing the Impact of Outliers on the One-Class Classification Decision Rule
A modified version of one-class classification criterion reducing the impact of outliers on the one-class classification decision rule is proposed based on support vector data description (SVDD) by D. Tax. The optimization method utilizes the substitution of nondifferentiable objective function by the smooth one. A comparative experimental study of existing one-class methods shows the superiority of the proposed criterion in anomaly detection.
Read moreVerification-Based Design of a Robust EMG Wake Word.
Surface electromyography (sEMG) signals are now commonly used in continuous myoelectric control of prostheses. More recently, researchers have considered EMG-based gesture recognition systems for human computer interaction research. These systems instead focus on recognizing discrete gestures (like a finger snap). The majority of works, however, have focused on improving multi-class performance, with little consideration for false activations from "other" classes. Consequently, they lack the robustness needed for real-world applications which generally require a single motion class such as a mouse click or a wake word. Furthermore, many works have borrowed the windowed classification schemes from continuous control, and thus fail to leverage the temporal structure of the gesture. In this paper, we propose a verification-based approach to creating a robust EMG wake word using one-class classifiers (Support Vector Data Description, One Class-Support Vector Machine, Dynamic Time Warping (DTW) & Hidden Markov Models). The area under the ROC curve (AUC) is used as a feature optimization objective as it provides a better representation of the verification performance. Equal error rate (EER) and AUC are then used as evaluation metrics. The results are computed using both window-based and temporal classifiers on a dataset consisting of five different gestures, with a best EER of 0.04 and AUC of 0.98, recorded using a DTW scheme. These results demonstrate a design framework that may benefit the development of more robust solutions for EMG-based wake words or input commands for a variety of interactive applications.
Read moreA Data-Driven Operating Performance Assessment Method based on Weighted Multi-Sphere Support Vector Data Description
In modern hot rolling process, operating performance assessment is of great practical significance for guiding the production adjustment for operators. From the perspective of classification, operating performance assessment is a multi-class classification problem. Since support vector data description (SVDD) is a one-class classifier, conventional methods usually construct an independent SVDD model for each class, which ignores the correlation among different classes. The hyperspheres of different classes may not be isolated but overlapped. If a test sample exists in the overlapping region, how to determine which class it belongs to is a knotty problem. Moreover, conventional methods treat all samples equally, but in practice, the sample number of different classes can be imbalanced, which will affect the classification performance of SVDD. In this study, an operating performance assessment method based on weighted multi-sphere SVDD (WMSVDD) is proposed for solving the aforementioned issues. WMSVDD considers the interactions among different classes in a unified way, optimizes the hyperspheres of different classes globally, and introduces a weight coefficient to the model for eliminating the affects of uneven class sizes. Simulation results on a real hot rolling process illustrate the effectiveness of the proposed method comparing to the traditional multi-class SVDD.
Read moreSolving one-class problem with outlier examples by SVM
Solving one-class problem with outlier examples by SVM
Monitoring of solid-state fermentation of wheat straw in a pilot scale using FT-NIR spectroscopy and support vector data description
Monitoring of solid-state fermentation of wheat straw in a pilot scale using FT-NIR spectroscopy and support vector data description
Read moreRisk Assessment Model Based on SVDD and Fuzzy Regression Method
The paper aims to solve the problem of insufficient high risk data in risk assessment of R&D projects. A one-class classification method called support vector data description (SVDD) is studied, and an intelligent risk assessment model based on SVDD with fuzzy regression information is also proposed. The model comes into being a new approach. Applying this approach, firstly verify the conversional risk evaluation indexes by fuzzy regression technique to develop a sensitive index system. Secondly the study uses the historical risk data referring to these indexes to train the SVDD one-class classifier. Unlike previously proposed intelligent methods of risk assessment, with this model the risk level can be distinguished only by training of low risk data. The results of its application on an example show that the method is feasible for risk assessment with the fuzzy high risk data.
Read more