- Research Article
158
- 10.1016/j.neucom.2016.05.081
Graph self-representation method for unsupervised feature selection
- Jun 24, 2016
- Neurocomputing
- Rongyao Hu + 6 more +6
Graph self-representation method for unsupervised feature selection
Selecting the informative features from the high dimensional data can improve the performance of the classification and get a deep understanding of the problems. A non-problem related feature contains little information and has little influence on the data distribution. By permuting the feature and calculating the data distribution difference, how much information the feature contains could be measured. In this paper, we propose an unsupervised feature selection method (EUFSPR), which combines the ensemble technique, clustering, permutation and data distribution evaluation techniques to measure the feature importance. Clustering is adopted to get the sample groups and the data distribution is evaluated by the overlapping areas. Eight gene expression microarray datasets are utilized to demonstrate the effectiveness of the proposed method over the unsupervised feature selection methods and supervised feature selection methods.
Graph self-representation method for unsupervised feature selection
Graph self-representation method for unsupervised feature selection
Simultaneous positive sequential vectors modeling and unsupervised feature selection via continuous hidden Markov models
Simultaneous positive sequential vectors modeling and unsupervised feature selection via continuous hidden Markov models
Balanced Spectral Feature Selection
In many real-world unsupervised learning applications, given data with balanced distribution, that is, there are an approximately equal number of instances in each class, we often need to construct a model to reveal such balance. However, in many data, especially the high-dimensional ones, the data in the original feature space often do not present such balance due to the redundant and noisy features. To tackle this problem, we apply an unsupervised spectral feature selection method to select some informative features, which can better reveal the balanced structure of data. Although spectral feature selection is one of the most popular unsupervised feature selection methods and has been widely studied, none of the existing spectral feature selection methods consider the balance property of data. To address this issue, in this article, we propose a novel balanced spectral feature selection (BSFS) method, which not only selects the discriminative features but also picks those to reveal the balanced structure of data. To the best of our knowledge, this is the first spectral feature selection method considering balance structure of data. By introducing a balanced regularization term, we integrate the balanced spectral clustering and feature selection into a unified framework seamlessly. At last, the experiments on benchmark datasets show that the proposed one outperforms the conventional feature selection methods in both clustering performance and balance, which demonstrates the effectiveness and efficiency of the proposed method.
Read moreWeak Monotonicity With Trend Analysis for Unsupervised Feature Evaluation.
Performance in an engineering system tends to degrade over time due to a variety of wearing or ageing processes. In supervisory controlled processes there are typically many signals being monitored that may help to characterize performance degradation. It is preferred to select the least amount of information to obtain high quality of predictive analysis from a large amount of collected data, in which labeling the data is not always feasible. To this end a novel unsupervised feature selection method, robust with respect to significant measurement disturbances, is proposed using the notion of ``weak monotonicity'' (WM). The robustness of this notion makes it very attractive to identify the common trend in the presence of measurement noises and population variation from the collected data. Based on WM, a novel suitability indicator is proposed to evaluate the performance of each feature. This new indicator is then used to select the key features that contribute to the WM of a family of processes when noises and variations among processes exist. In order to evaluate the performance of the proposed framework of the WM and suitability, a comparative study with other nine state-of-the-arts unsupervised feature evaluation and selection methods is carried out on well-known benchmark datasets. The results show a promising performance of the proposed framework on unsupervised feature evaluation in the presence of measurement noises and population variations.
Read moreUnsupervised feature selection based on variance–covariance subspace distance
Subspace distance is an invaluable tool exploited in a wide range of feature selection methods. The power of subspace distance is that it can identify a representative subspace, including a group of features that can efficiently approximate the space of original features. On the other hand, employing intrinsic statistical information of data can play a significant role in a feature selection process. Nevertheless, most of the existing feature selection methods founded on the subspace distance are limited in properly fulfilling this objective. To pursue this void, we propose a framework that takes a subspace distance into account which is called “Variance–Covariance subspace distance”. The approach gains advantages from the correlation of information included in the features of data, thus determines all the feature subsets whose corresponding Variance–Covariance matrix has the minimum norm property. Consequently, a novel, yet efficient unsupervised feature selection framework is introduced based on the Variance–Covariance distance to handle both the dimensionality reduction and subspace learning tasks. The proposed framework has the ability to exclude those features that have the least variance from the original feature set. Moreover, an efficient update algorithm is provided along with its associated convergence analysis to solve the optimization side of the proposed approach. An extensive number of experiments on nine benchmark datasets are also conducted to assess the performance of our method from which the results demonstrate its superiority over a variety of state-of-the-art unsupervised feature selection methods. The source code is available at https://github.com/SaeedKarami/VCSDFS.
Read moreNon-convex regularized self-representation for unsupervised feature selection
Non-convex regularized self-representation for unsupervised feature selection
Feature Selection and Feature Stability Measurement Method for High-Dimensional Small Sample Data Based on Big Data Technology.
With the rapid development of artificial intelligence in recent years, the research on image processing, text mining, and genome informatics has gradually deepened, and the mining of large-scale databases has begun to receive more and more attention. The objects of data mining have also become more complex, and the data dimensions of mining objects have become higher and higher. Compared with the ultra-high data dimensions, the number of samples available for analysis is too small, resulting in the production of high-dimensional small sample data. High-dimensional small sample data will bring serious dimensional disasters to the mining process. Through feature selection, redundancy and noise features in high-dimensional small sample data can be effectively eliminated, avoiding dimensional disasters and improving the actual efficiency of mining algorithms. However, the existing feature selection methods emphasize the classification or clustering performance of the feature selection results and ignore the stability of the feature selection results, which will lead to unstable feature selection results, and it is difficult to obtain real and understandable features. Based on the traditional feature selection method, this paper proposes an ensemble feature selection method, Random Bits Forest Recursive Clustering Eliminate (RBF-RCE) feature selection method, combined with multiple sets of basic classifiers to carry out parallel learning and screen out the best feature classification results, optimizes the classification performance of traditional feature selection methods, and can also improve the stability of feature selection. Then, this paper analyzes the reasons for the instability of feature selection and introduces a feature selection stability measurement method, the Intersection Measurement (IM), to evaluate whether the feature selection process is stable. The effectiveness of the proposed method is verified by experiments on several groups of high-dimensional small sample data sets.
Read moreAn Unsupervised Feature Selection Method for Data-Driven Anomaly Detection Systems
Feature selection has been widely used as a pre-processing step that helps to optimise the performance of data-driven intrusion/anomaly detection systems in achieving their tasks. For example, when grouping the data into normal and outlier groups, the existence of redundant and non-representative features would reduce the accuracy of classifying the data points and would also increase the processing time. Therefore, feature selection is applied as a pre-processing step for anomaly detection systems in order to optimize their classification accuracy and running time. Most of the existing feature selection methods have limitations when dealing with high-dimensional data, as they search different subsets of features to find accurate representations of all features. Obviously, searching for different combinations of features is computationally very expensive, which makes existing work not efficient for high-dimensional data. The work carried out here, which relates to the design of a similaritybased unsupervised feature selection method for an efficient and accurate anomaly detection (UFSAD), tackles mainly the selection of reduced set of representative features from high-dimensional data without the data class labels. The selected features should improve the accuracy and performance of anomaly detection systems due to the elimination of redundant and non-representative features. The proposed UFSAD method extends the k-mean clustering algorithm to partition the features into k clusters based on a similarity measure (e.g. PCC - Pearson Correlation Coefficient, LSRE - Least Square Regression Error or MICI - Maximal Information Compression Index) in order to accurately partition the features. Then the proposed centroid-based feature selection method is used, where the feature with the closest similarity to its cluster centroid is selected as the representative feature while others are discarded. Extensive experimental work has shown that UFSAD can generate a reduced representative and non-redundant feature set that achieves good classification accuracy in comparison with well-known unsupervised features selection methods.
Read moreMCDM-EFS: A novel ensemble feature selection method for software defect prediction using multi-criteria decision making
Software defect prediction models are used for predicting high risk software components. Feature selection has significant impact on the prediction performance of the software defect prediction models since redundant and unimportant features make the prediction model more difficult to learn. Ensemble feature selection has recently emerged as a new methodology for enhancing feature selection performance. This paper proposes a new multi-criteria-decision-making (MCDM) based ensemble feature selection (EFS) method. This new method is termed as MCDM-EFS. The proposed method, MCDM-EFS, first generates the decision matrix signifying the feature’s importance score with respect to various existing feature selection methods. Next, the decision matrix is used as the input to well-known MCDM method TOPSIS for assigning a final rank to each feature. The proposed approach is validated by an experimental study for predicting software defects using two classifiers K-nearest neighbor (KNN) and naïve bayes (NB) over five open-source datasets. The predictive performance of the proposed approach is compared with existing feature selection algorithms. Two evaluation metrics – nMCC and G-measure are used to compare predictive performance. The experimental results show that the MCDM-EFS significantly improves the predictive performance of software defect prediction models against other feature selection methods in terms of nMCC as well as G-measure.
Read moreLow-rank structure preserving for unsupervised feature selection
Low-rank structure preserving for unsupervised feature selection
Unsupervised Feature Selection via Unified Trace Ratio Formulation and K-means Clustering (TRACK)
Feature selection plays a crucial role in scientific research and practical applications. In the real world applications, labeling data is time and labor consuming. Thus, unsupervised feature selection methods are desired for many practical applications. Linear discriminant analysis (LDA) with trace ratio criterion is a supervised dimensionality reduction method that has shown good performance to improve classifications. In this paper, we first propose a unified objective to seamlessly accommodate trace ratio formulation and K-means clustering procedure, such that the trace ratio criterion is extended to unsupervised model. After that, we propose a novel unsupervised feature selection method by integrating unsupervised trace ratio formulation and structured sparsity-inducing norms regularization. The proposed method can harness the discriminant power of trace ratio criterion, thus it tends to select discriminative features. Meanwhile, we also provide two important theorems to guarantee the unsupervised feature selection process. Empirical results on four benchmark data sets show that the proposed method outperforms other sate-of-the-art unsupervised feature selection algorithms in all three clustering evaluation metrics.
Read moreUnsupervised feature selection based on the measures of degree of dependency using rough set theory in digital mammogram image classification
Feature Selection (FS) has become one of the most active research topics in the area of data mining. It performs to remove redundant and noisy features from high-dimensional data sets. A good feature selection has several advantages for a learning algorithm such as reducing computational cost, increasing its classification accuracy and improving result comprehensibility. In the supervised FS methods various feature subsets are evaluated using an evaluation function or metric to select only those features which are related to the decision classes of the data under consideration. However, for many data mining applications, decision class labels are often unknown or incomplete, thus indicating the significance of unsupervised feature selection. However, in unsupervised learning, decision class labels are not provided. The problem is that not all features are important. Some of the features may be redundant, and others may be irrelevant and noisy. In this paper, a novel unsupervised feature selection in mammogram image, using rough set based measures, is proposed. A typical mammogram image processing system generally consists of mammogram image acquisition, preprocessing of image, segmentation, features extracted from the segmented mammogram image. The proposed method is used to select features from data set, the method is compared with existing rough set based supervised feature selection methods and classification performance of both methods are recorded and demonstrates the efficiency of the method.
Read moreDependence Guided Unsupervised Feature Selection
In the past decade, various sparse learning based unsupervised feature selection methods have been developed. However, most existing studies adopt a two-step strategy, i.e., selecting the top-m features according to a calculated descending order and then performing K-means clustering, resulting in a group of sub-optimal features. To address this problem, we propose a Dependence Guided Unsupervised Feature Selection (DGUFS) method to select features and partition data in a joint manner. Our proposed method enhances the inter-dependence among original data, cluster labels, and selected features. In particular, a projection-free feature selection model is proposed based on l20-norm equality constraints. We utilize the learned cluster labels to fill in the information gap between original data and selected features. Two dependence guided terms are consequently proposed for our model. More specifically, one term increases the dependence of desired cluster labels on original data, while the other term maximizes the dependence of selected features on cluster labels to guide the process of feature selection. Last but not least, an iterative algorithm based on Alternating Direction Method of Multipliers (ADMM) is designed to solve the constrained minimization problem efficiently. Extensive experiments on different datasets consistently demonstrate that our proposed method significantly outperforms state-of-the-art baselines.
Read moreCompactness score: a fast filter method for unsupervised feature selection
The rapid development of big data era incurs the generation of huge amount of data day by day in various fields. Due to the large-scale and high-dimensional characteristics of these data, it is often difficult to achieve better decision-making in practical applications. Therefore, an efficient big data analytical method is urgently necessary. For feature engineering, feature selection seems to be an important research topic which is anticipated to select “excellent” features from candidate ones. The implementation of feature selection can not only achieve the purpose of dimensionality reduction, but also improve the computational efficiency and result performance of the model. In many classification tasks, researchers found that data seem to be usually close to each other if they are from the same class; thus, local compactness is of great importance for the evaluation of a feature. Based on this discovery, we propose a fast unsupervised feature selection algorithm, named Compactness Score (CSUFS), to select desired features. To prove the superiority of the proposed algorithm, several public data sets are considered with extensive experiments being performed. The experiments are presented by applying feature subsets selected through several different algorithms to the clustering task. The performance of clustering tasks is indicated by two well-known evaluation metrics, while the efficiency is reflected by the corresponding running time. As demonstrated, our proposed algorithm is more accurate and efficient compared with existing ones.
Read moreUnsupervised feature selection using binary bat algorithm
Feature selection is selecting a subset of optimal features. Feature selection is being used in high dimensional data reduction and it is being used in several applications like medical, image processing, text mining, etc. Several methods were introduced for unsupervised feature selection. Among those methods some are based on filter approach and some are based on wrapper approach. In the existing work, unsupervised feature selection methods using Genetic Algorithm, Particle Swarm Optimization with Relative Reduct, Quick Reduct and Ant Colony Optimization have been introduced. These methods yield better performance for unsupervised feature selection. In this paper we proposed a novel method to select subset of features from unlabeled data using binary bat algorithm with sum of squared error as the fitness function. The proposed method is then tested with various classification algorithms like decision tree, multilayer perceptron, support vector machine and clustering quality measures like sum of squared error. The results show that our proposed method gives more accuracy when compared with other optimization algorithm.
Read more