- Research Article
5
- 10.1016/j.asoc.2015.08.038
On learning dual classifiers for better data classification
- Aug 29, 2015
- Applied Soft Computing
- Wei-Chao Lin + 3 more +3
On learning dual classifiers for better data classification
Instance selection in medical datasets: A divide-and-conquer framework
On learning dual classifiers for better data classification
On learning dual classifiers for better data classification
A Novel Prediction Approach for Effective Medical Data Mining
Data mining techniques have been employed for solving many medical problems, especially for disease prediction. For instance, given a dataset containing normal and cancerous patients, the goal is to develop a model to predict whether a new (unknown) patient belongs to the normal or cancerous class. In general, the model is constructed based on some machine learning technique over a collected training set. However, the quality of the training set can affect the final prediction performance of the model. That is, if the training set contains some certain amount of noisy data (or outliers), then the model's performance could be degraded. In literature, instance selection is performed over a given training set in order to filter out some noisy data and the reduced training set containing non-noisy data is used for developing the prediction model. In this paper, we present a novel approach where instance selection is performed to divide a given training set into noisy and non-noisy subsets. Then, they are used to train two models respectively. During prediction, the instance selection step is also executed over the testing set, in which the noisy and non-noisy subsets are used to test their corresponding models respectively. The experimental results based on various medical domain datasets show that our proposed approach performs better than the baseline, which is based on the conventional instance selection approach.
Read moreCluster-oriented instance selection for classification problems
Cluster-oriented instance selection for classification problems
Large scale instance selection by means of federal instance selection
Large scale instance selection by means of federal instance selection
SVOIS: Support Vector Oriented Instance Selection for text classification
SVOIS: Support Vector Oriented Instance Selection for text classification
A Cooperative Coevolutionary Algorithm For KNN Training Set Optimization
The traditional evolutionary instance selection algorithm has the risk of redundant and noise training samples in the training set selection, which affects the classification effect. In this paper, instance selection, instance weighting and feature weighting are integrated into the cooperative coevolution framework, and a cooperative coevolutionary algorithm for KNN training set optimization selection is proposed. The CHC algorithm based on multi-point crossover strategy is used to further improve the accuracy of instance selection. The SSGA algorithm based on fast mutation strategy of instance weighting and feature weighting is synergistic with the instance selection, and it helps to remove noise and irrelevant data in the process of instance selection. The proposed method can speed up the convergence of the population, it can also improve the efficiency of the algorithm, and improve the KNN classification performance. The experimental results show that this method has advantages in classification accuracy and efficiency compared with some current evolutionary instance selection algorithms.
Read moreENRICHing medical imaging training sets enables more efficient machine learning.
Deep learning (DL) has been applied in proofs of concept across biomedical imaging, including across modalities and medical specialties. Labeled data are critical to training and testing DL models, but human expert labelers are limited. In addition, DL traditionally requires copious training data, which is computationally expensive to process and iterate over. Consequently, it is useful to prioritize using those images that are most likely to improve a model's performance, a practice known as instance selection. The challenge is determining how best to prioritize. It is natural to prefer straightforward, robust, quantitative metrics as the basis for prioritization for instance selection. However, in current practice, such metrics are not tailored to, and almost never used for, image datasets. To address this problem, we introduce ENRICH-Eliminate Noise and Redundancy for Imaging Challenges-a customizable method that prioritizes images based on how much diversity each image adds to the training set. First, we show that medical datasets are special in that in general each image adds less diversity than in nonmedical datasets. Next, we demonstrate that ENRICH achieves nearly maximal performance on classification and segmentation tasks on several medical image datasets using only a fraction of the available images and without up-front data labeling. ENRICH outperforms random image selection, the negative control. Finally, we show that ENRICH can also be used to identify errors and outliers in imaging datasets. ENRICH is a simple, computationally efficient method for prioritizing images for expert labeling and use in DL.
Read moreBagging of Instance Selection Algorithms
The paper presents bagging ensembles of instance selection algorithms. We use bagging to improve instance selection. The improvement comprises data compression and prediction accuracy. The examined instance selection algorithms for classification are ENN, CNN, RNG and GE and for regression are the developed by us Generalized CNN and Generalized ENN algorithms. Results of the comparative experimental study performed using different configurations on several datasets shows that the approachbased on bagging allowed for significant improvement, especially in terms of data compression.
Read moreImproving Instance Selection via Metric Learning
The k-Nearest Neighbor (k-NN) rule is widely used for classification tasks because of its simplicity and efficiency. However, a well-known drawback of k-NN is its dependence on the quality of the training set, since the k-NN makes no assumption about the importance of each instance. In fact, the existence of noisy and superfluous instances in the training set tends to increase the classification error rate. Thus, instance selection methods are useful to identify which instances belonging to the training set will be considered in the k-NN classifier. Our proposal shows a simple and effective way to improve instance selection methods using metric learning. The idea of our proposal relies on a pure geometric intuition that metric learning transforms the input space where points in the same class are simultaneously near each other and far from points in the other classes. In a more “organised” space, we show that instance selection methods can benefit from this transformed space. We carried out an experimental evaluation to compare the instance selection with and without metric learning on UCI benchmark data sets. The results reveals that the combination of metric and instance selection is very welcome. All tested instance selection methods improved significantly.
Read moreCluster-based instance selection for machine classification
Instance selection in the supervised machine learning, often referred to as the data reduction, aims at deciding which instances from the training set should be retained for further use during the learning process. Instance selection can result in increased capabilities and generalization properties of the learning model, shorter time of the learning process, or it can help in scaling up to large data sources. The paper proposes a cluster-based instance selection approach with the learning process executed by the team of agents and discusses its four variants. The basic assumption is that instance selection is carried out after the training data have been grouped into clusters. To validate the proposed approach and to investigate the influence of the clustering method used on the quality of the classification, the computational experiment has been carried out.
Read moreInstance Selection Optimization for Neural Network Training
Performing instance selection prior to the classifier training is always beneficial in terms of computational complexity reduction of the classifier training and sometimes also beneficial in terms of improving prediction accuracy. Removing the noisy instances improves the prediction accuracy and removing redundant and irrelevant instances does not negatively effect it. However, in practice the instance selection methods usually also remove some instances, which should not be removed from the training dataset, what results in decreasing the prediction accuracy. We discuss two methods to deal with the problem. The first method is the parameterization of instance selection algorithms, which allows to choose how aggressively the instances are removed and the second one is to embed the instance selection directly into the prediction model, which in our case is an MLP neural network.
Read moreA review of instance selection methods
In supervised learning, a training set providing previously known information is used to classify new instances. Commonly, several instances are stored in the training set but some of them are not useful for classifying therefore it is possible to get acceptable classification rates ignoring non useful cases; this process is known as instance selection. Through instance selection the training set is reduced which allows reducing runtimes in the classification and/or training stages of classifiers. This work is focused on presenting a survey of the main instance selection methods reported in the literature.
Read moreComparison of Instance Selection Algorithms II. Results and Comments
This paper is an continuation of the accompanying paper with the same main title. The first paper reviewed instance selection algorithms, here results of empirical comparison and comments are presented. Several test were performed mostly on benchmark data sets from the machine learning repository at UCI. Instance selection algorithms were tested with neural networks and machine learning algorithms.
Read moreDIER ‐Net: Debiased Learning With Medical Image Noisy Label by Intrinsic and Extrinsic Regularization
In medical image analysis, the presence of noisy labels and imbalanced data poses significant challenges to the performance of deep learning models, particularly in critical diagnostic tasks. To address this issue, we propose DIER‐Net, a learning with noisy label framework designed to handle noisy labels in imbalanced medical datasets. Our approach introduces a debiased sample selection technique that effectively filters out noisy labels while preserving important minority class samples. Additionally, we employ intrinsic and extrinsic regularization strategies to enhance the model's robustness by leveraging both clean and noisy data. Our method is evaluated on two widely used medical image datasets: the ISIC melanoma classification and Kaggle histopathologic lymph node classification. The experimental results demonstrate that DIER‐Net consistently outperforms existing state‐of‐the‐art methods, particularly in settings with high levels of label noise, offering a robust solution for real‐world clinical applications where noisy and imbalanced data are common. DIER‐Net provides an effective approach to enhance the reliability of AI systems in medical imaging, contributing to more accurate and trustworthy diagnostic outcomes.
Read moreSelf-Configuring Hybrid Evolutionary Algorithm for Fuzzy Imbalanced Classification with Adaptive Instance Selection
A novel approach for instance selection in classification problems is presented. This adaptive instance selection is designed to simultaneously decrease the amount of computation resources required and increase the classification quality achieved. The approach generates new training samples during the evolutionary process and changes the training set for the algorithm. The instance selection is guided by means of changing probabilities, so that the algorithm concentrates on problematic examples which are difficult to classify. The hybrid fuzzy classification algorithm with a self-configuration procedure is used as a problem solver. The classification quality is tested upon 9 problem data sets from the KEEL repository. A special balancing strategy is used in the instance selection approach to improve the classification quality on imbalanced datasets. The results prove the usefulness of the proposed approach as compared with other classification methods.
Read more