- Research Article
40
- 10.1016/j.patcog.2014.08.019
Noise-robust semi-supervised learning via fast sparse coding
- Aug 28, 2014
- Pattern Recognition
- Zhiwu Lu + 1 more +1
Noise-robust semi-supervised learning via fast sparse coding
Graph-based semi-supervised learning plays an important role in large scale image classification tasks. However, the problem becomes very challenging in the presence of noisy labels and outliers. Moreover, traditional robust semi-supervised learning solutions suffers from prohibitive computational burdens thus cannot be computed for streaming data. Motivated by that, we present a novel unified framework robust structure-aware semi-supervised learning called Unified RSSL (URSSL) for batch processing and recursive processing robust to both outliers and noisy labels. Particularly, URSSL applies joint semi-supervised dimensionality reduction with robust estimators and network sparse regularization simultaneously on the graph Laplacian matrix iteratively to preserve the intrinsic graph structure and ensure robustness to the compound noise. First, in order to relieve the influence from outliers, a novel semi-supervised robust dimensionality reduction is applied relying on robust estimators to suppress outliers. Meanwhile, to tackle noisy labels, the denoised graph similarity information is encoded into the network regularization. Moreover, by identifying strong relevance of dimensionality reduction and network regularization in the context of robust semi-supervised learning (RSSL), a two-step alternative optimization is derived to compute optimal solutions with guaranteed convergence. We further derive our framework to adapt to large scale semi-supervised learning particularly suitable for large scale image classification and demonstrate the model robustness under different adversarial attacks. For recursive processing, we rely on reparameterization to transform the formulation to unlock the challenging problem of robust streaming-based semi-supervised learning. Last but not least, we extend our solution into distributed solutions to resolve the challenging issue of distributed robust semi-supervised learning when images are captured by multiple cameras at different locations. Extensive experimental results demonstrate the promising performance of this framework when applied to multiple benchmark datasets with respect to state-of-the-art approaches for important applications in the areas of image classification and spam data analysis.
Noise-robust semi-supervised learning via fast sparse coding
Noise-robust semi-supervised learning via fast sparse coding
ATRMME: A Tensor Ring-Based Feature Extraction Method for Hyperspectral Remote Sensing Image Classification Under Noisy Labels
The extraction of concise and discriminative features from hyperspectral remote sensing images (HSIs) holds great significance for land cover classification tasks. Nevertheless, the intricate spatial-spectral tensor structure and inaccurate class labels undermine the performance of existing vector-based or Tucker tensor-based feature dimensionality reduction (DR) approaches. To tackle this problem, we propose an Adaptive Tensor-Ring Manifold Maintaining Embedding (ATRMME) approach for HSI feature DR in the presence of noisy labels. Firstly, a tensor-ring subspace representation theory is established to characterize the low-dimensional hyperspectral tensor subspace. To mine concise features from complex spatial-spectral data, differing from existing Tucker-based DR methods that involve exorbitant parameter overheads, a nuclear norm-constrained Tensor-Ring (TR) feature representation term is designed. This term enables the compression of key features within the low-dimensional subspace with a linear parameter cost. For the adaptive perception and capture of the manifold structure under noisy labels, an adaptive adjacency matrix is constructed, and the adjacency relationships are reorganized based on the TR feature distribution. To guarantee the effective execution of model training, an alternating optimization algorithm is developed for parameter estimation. Experimental results on multiple hyperspectral datasets verify that the ATRMME approach exhibits strong robustness against noisy labels, outperforms conventional vector-based and Tucker tensor-based methods, and has lower parameter requirements.
Read moreSemi-supervised Dictionary Active Learning for Pattern Classification
Gathering labeled data is one of the most time-consuming and expensive tasks in supervised machine learning. In practical applications, there are usually quite limited labeled training samples but abundant unlabeled data that is easy to collect. Semi-supervised learning and active learning are two important techniques for learning a discriminative classification model when labeled data is scarce. However, unlabeled data with significant noises and outliers cannot be well exploited and usually worsen the performance of semi-supervised learning and the performance of active learning also needs a powerful initial classifier learned from the quite limited labeled training data. In order to solve the above issues, in this paper we proposed a novel model of semi-supervised dictionary active learning (SSDAL), which aims to integrate semi-supervised learning and active learning to effectively use all the training data. In particular, two criterions based on estimated class possibility are designed to select the unlabeled data with confident class estimation for semi-supervised learning and the informative unlabeled data for active learning, respectively. Extensive experiments are conducted to show the superior performance of our method in classification applications, e.g., handwritten digit recognition, face recognition and large-scale image classification.
Read moreHigh-dimensional semi-supervised learning via a fusion-refinement procedure
High-dimensional semi-supervised learning via a fusion-refinement procedure
NoiseBox: Toward More Efficient and Effective Learning With Noisy Labels
Despite the large progress in supervised learning with neural networks, there are significant challenges in obtaining high-quality, large-scale and accurately labelled datasets. In such contexts, how to learn in the presence of noisy labels has received more and more attention. Addressing this relatively intricate problem to attain competitive results predominantly involves designing mechanisms that select samples that are expected to have reliable annotations. However, these methods typically involve multiple off-the-shelf techniques, resulting in intricate structures. Furthermore, they frequently make implicit or explicit assumptions about the noise modes/ratios within the dataset. Such assumptions can compromise model robustness and limit its performance under varying noise conditions. Unlike these methods, in this work, we propose an efficient and effective framework with minimal hyperparameters that achieves SOTA results in various benchmarks. Specifically, we design an efficient and concise training framework consisting of a subset expansion module responsible for exploring non-selected samples and a model training module to further reduce the impact of noise, called NoiseBox. Moreover, diverging from common sample selection methods based on the “small loss” mechanism, we introduce a novel sample selection method based on the neighbouring relationships and label consistency in the feature space. Without bells and whistles, such as model co-training, self-supervised pre-training and semi-supervised learning, and with robustness concerning the settings of its few hyper-parameters, our method significantly surpasses previous methods on both CIFAR10/CIFAR100 with synthetic noise and real-world noisy datasets such as Red Mini-ImageNet, WebVision, Clothing1M and ANIMAL-10N.
Read moreSemi-Supervised Learning with Density-Sensitive Manifold graph
The key problem of Graph-Based Semi-Supervised Learning (GBSSL) methods is how to construct the graph structure under some assumptions. While distance information among graph nodes is investigated well for graph construction, the density information is not given enough attention. In this paper, we propose a novel GBSSL method, named Density-Sensitive Manifold Learning (DSML), which introduces density distribution into graph construction by calculating a new propagation coefficient matrix. The experimental results show that DSML scheme performs better than traditional GBSSL methods. More importantly, the new propagation coefficient matrix can be easily introduced into traditional GBSSL methods to improve their performance, which is also validated in the experiments.
Read moreConstrained graph-based semi-supervised learning with higher order regularization
Graph-based semi-supervised learning (SSL) algorithms have been widely studied in the last few years. Most of these algorithms were designed from unconstrained optimization problems using a Laplacian regularizer term as smoothness functional in an attempt to reflect the intrinsic geometric structure of the data's marginal distribution. Although a number of recent research papers are still focusing on unconstrained methods for graph-based SSL, a recent statistical analysis showed that many of these algorithms may be unstable on transductive regression. Therefore, we focus on providing new constrained methods for graph-based SSL. We begin by analyzing the regularization framework of existing unconstrained methods. Then, we incorporate two normalization constraints into the optimization problem of three of these methods. We show that the proposed optimization problems have closed-form solution. By generalizing one of these constraints to any distribution, we provide generalized methods for constrained graph-based SSL. The proposed methods have a more flexible regularization framework than the corresponding unconstrained methods. More precisely, our methods can deal with any graph Laplacian and use higher order regularization, which is effective on general SSL taks. In order to show the effectiveness of the proposed methods, we provide comprehensive experimental analyses. Specifically, our experiments are subdivided into two parts. In the first part, we evaluate existing graph-based SSL algorithms on time series data to find their weaknesses. In the second part, we evaluate the proposed constrained methods against six state-of-the-art graph-based SSL algorithms on benchmark data sets. Since the widely used best case analysis may hide useful information concerning the SSL algorithms' performance with respect to parameter selection, we used recently proposed empirical evaluation models to evaluate our results. Our results show that our methods outperforms the competing methods on most parameter settings and graph construction methods. However, we found a few experimental settings in which our methods showed poor performance. In order to facilitate the reproduction of our results, the source codes, data sets, and experimental results are freely available.
Read moreRobust semi-supervised learning in open environments
Semi-supervised learning (SSL) aims to improve performance by exploiting unlabeled data when labels are scarce. Conventional SSL studies typically assume close environments where important factors (e.g., label, feature, distribution) between labeled and unlabeled data are consistent. However, more practical tasks involve open environments where important factors between labeled and unlabeled data are inconsistent. It has been reported that exploiting inconsistent unlabeled data causes severe performance degradation, even worse than the simple supervised learning baseline. Manually verifying the quality of unlabeled data is not desirable, therefore, it is important to study robust SSL with inconsistent unlabeled data in open environments. This paper briefly introduces some advances in this line of research, focusing on techniques concerning label, feature, and data distribution inconsistency in SSL, and presents the evaluation benchmarks. Open research problems are also discussed for reference purposes.
Read moreDIER ‐Net: Debiased Learning With Medical Image Noisy Label by Intrinsic and Extrinsic Regularization
In medical image analysis, the presence of noisy labels and imbalanced data poses significant challenges to the performance of deep learning models, particularly in critical diagnostic tasks. To address this issue, we propose DIER‐Net, a learning with noisy label framework designed to handle noisy labels in imbalanced medical datasets. Our approach introduces a debiased sample selection technique that effectively filters out noisy labels while preserving important minority class samples. Additionally, we employ intrinsic and extrinsic regularization strategies to enhance the model's robustness by leveraging both clean and noisy data. Our method is evaluated on two widely used medical image datasets: the ISIC melanoma classification and Kaggle histopathologic lymph node classification. The experimental results demonstrate that DIER‐Net consistently outperforms existing state‐of‐the‐art methods, particularly in settings with high levels of label noise, offering a robust solution for real‐world clinical applications where noisy and imbalanced data are common. DIER‐Net provides an effective approach to enhance the reliability of AI systems in medical imaging, contributing to more accurate and trustworthy diagnostic outcomes.
Read moreCombining Smooth Graphs with Semi-supervised Learning
The key points of the semi-supervised learning problem are the label smoothness and cluster assumptions. In graph-based semi-supervised learning, graph representations of the data are so important that different graph representations can affect the classification results heavily. We present a novel method to produce a graph called smooth Markov random walk graph which takes into account the two assumptions employed by semi-supervised learning. The new graph is achieved by modifying the eigenspectrum of the transition matrix of Markov random walk graph and is sufficiently smooth with respect to the intrinsic structure of labeled and unlabeled points.We believe the smoother graph will benefit semi-supervised learning. Experiments on artificial and real world dataset indicate that our method provides superior classification accuracy over several state-of-the-art methods.
Read moreSpeaker Attribution with Voice Profiles by Graph-Based Semi-Supervised Learning
Speaker attribution is required in many real-world applications, such as\nmeeting transcription, where speaker identity is assigned to each utterance\naccording to speaker voice profiles. In this paper, we propose to solve the\nspeaker attribution problem by using graph-based semi-supervised learning\nmethods. A graph of speech segments is built for each session, on which\nsegments from voice profiles are represented by labeled nodes while segments\nfrom test utterances are unlabeled nodes. The weight of edges between nodes is\nevaluated by the similarities between the pretrained speaker embeddings of\nspeech segments. Speaker attribution then becomes a semi-supervised learning\nproblem on graphs, on which two graph-based methods are applied: label\npropagation (LP) and graph neural networks (GNNs). The proposed approaches are\nable to utilize the structural information of the graph to improve speaker\nattribution performance. Experimental results on real meeting data show that\nthe graph based approaches reduce speaker attribution error by up to 68%\ncompared to a baseline speaker identification approach that processes each\nutterance independently.\n
Read moreTowards Dynamic Self-Training for Scalable Semi-Supervised Learning on Graphs
Towards Dynamic Self-Training for Scalable Semi-Supervised Learning on Graphs
Semi-Supervised Local Fisher Discriminant Analysis Based on Reconstruction Probability Class
Fisher discriminant analysis (FDA) is a classic supervised dimensionality reduction method in statistical pattern recognition. FDA can maximize the scatter between different classes, while minimizing the scatter within each class. As it only utilizes the labeled data and ignores the unlabeled data in the analysis process of FDA, it cannot be used to solve the unsupervised learning problems. Its performance is also very poor in dealing with semi-supervised learning problems in some cases. Recently, several semi-supervised learning methods as an extension of FDA have proposed. Most of these methods solve the semi-supervised problem by using a tradeoff parameter that evaluates the ratio of the supervised and unsupervised methods. In this paper, we propose a general semi-supervised dimensionality learning idea for the partially labeled data, namely the reconstruction probability class of labeled and unlabeled data. Based on the probability class optimizes Fisher criterion function, we propose a novel Semi-Supervised Local Fisher Discriminant Analysis (S2LFDA) method. Experimental results on real-world datasets demonstrate its effectiveness compared to the existing similar correlation methods.
Read moreSemi-supervised and Active Learning Models for Software Fault Prediction
As software continues to insinuate itself into nearly every aspect of our life, the quality of software has been an extremely important issue. Software Quality Assurance (SQA) is a process that ensures the development of high-quality software. It concerns the important problem of maintaining, monitoring, and developing quality software. Accurate detection of fault prone components in software projects is one of the most commonly practiced techniques that offer the path to high quality products without excessive assurance expenditures. This type of quality modeling requires the availability of software modules with known fault content developed in similar environment. However, collection of fault data at module level, particularly in new projects, is expensive and time-consuming. Semi-supervised learning and active learning offer solutions to this problem for learning from limited labeled data by utilizing inexpensive unlabeled data.;In this dissertation, we investigate semi-supervised learning and active learning approaches in the software fault prediction problem. The role of base learner in semi-supervised learning is discussed using several state-of-the-art supervised learners. Our results showed that semi-supervised learning with appropriate base learner leads to better performance in fault proneness prediction compared to supervised learning. In addition, incorporating pre-processing technique prior to semi-supervised learning provides a promising direction to further improving the prediction performance. Active learning, sharing the similar idea as semi-supervised learning in utilizing unlabeled data, requires human efforts for labeling fault proneness in its learning process. Empirical results showed that active learning supplemented by dimensionality reduction technique performs better than the supervised learning on release-based data sets.
Read moreKnowledge Distillation Meets Label Noise Learning: Ambiguity-Guided Mutual Label Refinery.
Knowledge distillation (KD), which aims at transferring the knowledge from a complex network (a teacher) to a simpler and smaller network (a student), has received considerable attention in recent years. Typically, most existing KD methods work on well-labeled data. Unfortunately, real-world data often inevitably involve noisy labels, thus leading to performance deterioration of these methods. In this article, we study a little-explored but important issue, i.e., KD with noisy labels. To this end, we propose a novel KD method, called ambiguity-guided mutual label refinery KD (AML-KD), to train the student model in the presence of noisy labels. Specifically, based on the pretrained teacher model, a two-stage label refinery framework is innovatively introduced to refine labels gradually. In the first stage, we perform label propagation (LP) with small-loss selection guided by the teacher model, improving the learning capability of the student model. In the second stage, we perform mutual LP between the teacher and student models in a mutual-benefit way. During the label refinery, an ambiguity-aware weight estimation (AWE) module is developed to address the problem of ambiguous samples, avoiding overfitting these samples. One distinct advantage of AML-KD is that it is capable of learning a high-accuracy and low-cost student model with label noise. The experimental results on synthetic and real-world noisy datasets show the effectiveness of our AML-KD against state-of-the-art KD methods and label noise learning (LNL) methods. Code is available at https://github.com/Runqing-forMost/ AML-KD.
Read more