- Research Article
15
- 10.1016/j.knosys.2022.109093
Partial multi-label learning via specific label disambiguation
- May 26, 2022
- Knowledge-Based Systems
- Feng Li + 2 more +2
Partial multi-label learning via specific label disambiguation
Partial multi-label learning (PML) models the scenario where each training instance is annotated with a set of candidate labels, and only some of the labels are relevant. The PML problem is practical in real-world scenarios, as it is difficult and even impossible to obtain precisely labeled samples. Several PML solutions have been proposed to combat with the prone misled by the irrelevant labels concealed in the candidate labels, but they generally focus on the smoothness assumption in feature space or low-rank assumption in label space, while ignore the negative information between features and labels. Specifically, if two instances have largely overlapped candidate labels, irrespective of their feature similarity, their ground-truth labels should be similar; while if they are dissimilar in the feature and candidate label space, their ground-truth labels should be dissimilar with each other. To achieve a credible predictor on PML data, we propose a novel approach called PML-LFC (Partial Multi-label Learning with Label and Feature Collaboration). PML-LFC estimates the confidence values of relevant labels for each instance using the similarity from both the label and feature spaces, and trains the desired predictor with the estimated confidence values. PML-LFC achieves the predictor and the latent label matrix in a reciprocal reinforce manner by a unified model, and develops an alternative optimization procedure to optimize them. Extensive empirical study on both synthetic and real-world datasets demonstrates the superiority of PML-LFC.
Partial multi-label learning via specific label disambiguation
Partial multi-label learning via specific label disambiguation
Prior Knowledge Regularized Self-Representation Model for Partial Multilabel Learning.
Partial multilabel learning (PML) aims to learn from training data, where each instance is associated with a set of candidate labels, among which only a part is correct. The common strategy to deal with such a problem is disambiguation, that is, identifying the ground-truth labels from the given candidate labels. However, the existing PML approaches always focus on leveraging the instance relationship to disambiguate the given noisy label space, while the potentially useful information in label space is not effectively explored. Meanwhile, the existence of noise and outliers in training data also makes the disambiguation operation less reliable, which inevitably decreases the robustness of the learned model. In this article, we propose a prior label knowledge regularized self-representation PML approach, called PAKS, where the self-representation scheme and prior label knowledge are jointly incorporated into a unified framework. Specifically, we introduce a self-representation model with a low-rank constraint, which aims to learn the subspace representations of distinct instances and explore the high-order underlying correlation among different instances. Meanwhile, we incorporate prior label knowledge into the above self-representation model, where the prior label knowledge is regarded as the complement of features to obtain an accurate self-representation matrix. The core of PAKS is to take advantage of the data membership preference, which is derived from the prior label knowledge, to purify the discovered membership of the data and accordingly obtain more representative feature subspace for model induction. Enormous experiments on both synthetic and real-world datasets show that our proposed approach can achieve superior or comparable performance to state-of-the-art approaches.
Read moreNoisy label tolerance: A new perspective of Partial Multi-Label Learning
Noisy label tolerance: A new perspective of Partial Multi-Label Learning
Partial multi-label feature selection based on label disambiguation and double-regularized sparse regression
Partial multi-label feature selection based on label disambiguation and double-regularized sparse regression
Partial Multi-Label Learning with Noisy Label Identification
Partial multi-label learning (PML) deals with problems where each instance is assigned with a candidate label set, which contains multiple relevant labels and some noisy labels. Recent studies usually solve PML problems with the disambiguation strategy, which recovers ground-truth labels from the candidate label set by simply assuming that the noisy labels are generated randomly. In real applications, however, noisy labels are usually caused by some ambiguous contents of the example. Based on this observation, we propose a partial multi-label learning approach to simultaneously recover the ground-truth information and identify the noisy labels. The two objectives are formalized in a unified framework with trace norm and ℓ1 norm regularizers. Under the supervision of the observed noise-corrupted label matrix, the multi-label classifier and noisy label identifier are jointly optimized by incorporating the label correlation exploitation and feature-induced noise model. Extensive experiments on synthetic as well as real-world data sets validate the effectiveness of the proposed approach.
Read moreHierarchical GAN-Tree and Bi-Directional Capsules for multi-label image classification
Hierarchical GAN-Tree and Bi-Directional Capsules for multi-label image classification
Multi-Label Learning via Feature and Label Space Dimension Reduction
In multi-label learning, each object belongs to multiple class labels simultaneously. In the data explosion age, the size of data is often huge, i.e., large number of instances, features and class labels. The high dimension of both the feature and label spaces has posed great challenges to multi-label learning problems, e.g., high time and memory costs. In this paper, we propose a new framework for multi-label learning with a large number of class labels and features, i.e., M ulti- L abel L earning via F eature and L abel S pace D imension R eduction, namely MLL-FLSDR. Specifically, both the feature space and label space are reduced to low dimensional spaces respectively, in which the local structure of data points is utilized to constrain the geometrical structure on both the learned low dimensional spaces and guarantee the qualities of them. Then, an effective multi-label classifier is constructed from the low dimensional feature space to the latent label space. Last, the final prediction for new test data examples can be obtained by recovering from their prediction results in the latent label space with an encoding matrix learned in the previous stage. Extensive comparison experiments with the state-of-the-art approaches manifest the effectiveness of the proposed method MLL-FLSDR.
Read morePartial Label Clustering
Partial label learning (PLL) is a significant weakly supervised learning framework, where each training example corresponds to a set of candidate labels and only one label is the ground-truth label. For the first time, this paper investigates the partial label clustering problem, which takes advantage of the limited available partial labels to improve the clustering performance. Specifically, we first construct a weight matrix of examples based on their relationships in the feature space and disambiguate the candidate labels to estimate the ground-truth label based on the weight matrix. Then, we construct a set of must-link and cannot-link constraints based on the disambiguation results. Moreover, we propagate the initial must-link and cannot-link constraints based on an adversarial prior promoted dual-graph learning approach. Finally, we integrate weight matrix construction, label disambiguation, and pairwise constraints propagation into a joint model to achieve mutual enhancement. We also theoretically prove that a better disambiguated label matrix can help improve clustering performance. Comprehensive experiments demonstrate our method realizes superior performance when comparing with state-of-the-art constrained clustering methods, and outperforms PLL and semi-supervised PLL methods when only limited samples are annotated. The code and appendix are publicly available at https://github.com/xyt-ml/PLC.
Read morePartial Label Dimensionality Reduction via Confidence-Based Dependence Maximization
Partial label learning deals with training examples each associated with a set of candidate labels, among which only one is valid. Most existing works focus on manipulating the label space by estimating the labeling confidences of candidate labels, while the task of manipulating the feature space by dimensionality reduction has been rarely investigated. In this paper, a novel partial label dimensionality reduction approach named CENDA is proposed via confidence-based dependence maximization. Specifically, CENDA adapts the Hilbert-Schmidt Independence Criterion (HSIC) to help identify the projection matrix, where the dependence between projected feature information and confidence-based labeling information is maximized iteratively. In each iteration, the projection matrix admits closed-form solution by solving a tailored generalized eigenvalue problem, while the labeling confidences of candidate labels are updated by conducting kNN aggregation in the projected feature space. Extensive experiments over a broad range of benchmark data sets show that the predictive performance of well-established partial label learning algorithms can be significantly improved by coupling with the proposed dimensionality reduction approach.
Read moreDeep Graph Matching for Partial Label Learning
Partial Label Learning (PLL) aims to learn from training data where each instance is associated with a set of candidate labels, among which only one is correct. In this paper, we formulate the task of PLL problem as an ``instance-label'' matching selection problem, and propose a DeepGNN-based graph matching PLL approach to solve it. Specifically, we first construct all instances and labels as graph nodes into two different graphs respectively, and then integrate them into a unified matching graph by connecting each instance to its candidate labels. Afterwards, the graph attention mechanism is adopted to aggregate and update all nodes state on the instance graph to form structural representations for each instance. Finally, each candidate label is embedded into its corresponding instance and derives a matching affinity score for each instance-label correspondence with a progressive cross-entropy loss. Extensive experiments on various data sets have demonstrated the superiority of our proposed method.
Read moreImproving Multi-Scenario Learning to Rank in E-commerce by Exploiting Task Relationships in the Label Space
Traditional Learning to Rank (LTR) models in E-commerce are usually trained on logged data from a single domain. However, data may come from multiple domains, such as hundreds of countries in international E-commerce platforms. Learning a single ranking function obscures domain differences, while learning multiple functions for each domain may also be inferior due to ignoring the correlations between domains. It can be formulated as a multi-task learning problem where multiple tasks share the same feature and label space. To solve the above problem, which we name Multi-Scenario Learning to Rank, we propose the Hybrid of implicit and explicit Mixture-of-Experts (HMoE) approach. Our proposed solution takes advantage of Multi-task Mixture-of-Experts to implicitly identify distinctions and commonalities between tasks in the feature space, and improves the performance with a stacked model learning task relationships in the label space explicitly. Furthermore, to enhance the flexibility, we propose an end-to-end optimization method with a task-constrained back-propagation strategy. We empirically verify that the optimization method is more effective than two-stage optimization required by the stacked approach. Experiments on real-world industrial datasets demonstrate that HMoE significantly outperforms the popular multi-task learning methods. HMoE is in-use in the search system of AliExpress and achieved 1.92% revenue gain in the period of one-week online A/B testing. We also release a sampled version of our dataset to facilitate future research.
Read moreMulti-Label Learning With Label Specific Features Using Correlation Information
To deal with the problem where each instance is associated with multiple labels, a lot of multi-label learning algorithms have been developed in recent years. Some approaches have been proposed to select label-specific features to utilize discriminate features for multi-label classification. Although label correlation has been considered in learning label-specific features, the critical correlation among instances was less taken into account. In this paper, we proposed a new approach called multi-label learning with label-specific features using correlation information (LSF-CI) to learn label-specific features for each label with the consideration of both correlation information in label space and correlation information in feature space. In the LSF-CI, the instance correlation in feature space is computed by a probabilistic neighborhood graph model, and label correlation in label space is computed by cosine similarity. For multi-label data, the LSF-CI has the capability to select Label-specific features for each label as well as classify an unseen instance into a set of relevant labels. To validate the effectiveness of LSF-CI, we conducted comprehensive experiments on eight multi-label datasets. The experimental results demonstrate that the LSF-CI is capable of selecting compact label-specific features, and achieving a competitive performance in comparison with the performances of the existing multi-label learning approaches.
Read morePartial Label Learning by Entropy Minimization
Partial label learning deals with the problem where each training example is associated with a set of candidate labels, only one of which is assumed to be valid. To learn from such ambiguous labeling information, the critical point is to disambiguate the set of candidate labels, thereby targeting the ground-truth label. By utilizing the nature that only one of the candidate labels is correct, we employ the entropy minimization strategy to force the model making confident predictions of the training data. By doing this, the ground-truth labels are likely to make more contributions to the model training. Finally, comparative experiments on a number of real-world datasets are conducted, demonstrating the effectiveness of the proposed approach.
Read moreA Generative Model for Partial Label Learning
Partial label (PL) learning tackles the problem where each training instance is associated with a set of candidate labels, among which only one is the true label. In this paper, we propose a novel generative model PL-CGAN, which tackles the partial label learning problem with Conditional Generative Adversarial Nets (CGAN). Specially, PL-CGAN introduces a teacher model to refine a more reliable soft label vector of each training instance by iteratively ensembling the current learned prediction network with the formal one in an online manner. Besides, it adopts a MixUp data augmentation scheme to prevent the prediction network from overfitting to the noisy labels. In addition, it deploys a CGAN to generate training instances with its corresponding label vectors, which form a feature alignment based on consistency cost to enhance it’s label refinement capacity. Extensive experiments are conducted on synthesized and real-world partial label learning datasets, while the proposed approach demonstrates the state-of-the-art performance for partial label learning.
Read moreSet Operation Aided Network for Action Units Detection
As a large number of parameters exist in deepmodel based methods, training such models usually requires many fully AU-annotated facial images. This is true with regard to the number of frames in two widely used datasets: BP4D[31] and DISFA [18], while those frames were captured from a small number of subjects (41, 27 respectively). This is problematic, as subjects produce highly consistent facial muscle movements, adding more frames per subject would only adds more close points in the feature space, and thus the classifier does not benefit from those extra frames. Data augmentation methods can be applied to alleviate the problem to a certain degree, but they fail to augment new subjects. We propose a novel Set Operation Aided Network (SO-Net) for action units detection. Specifically, new features and the corresponding labels are generated by adding set operations to both the feature and label spaces. The generated new features can be treated as a representation of a hypothetical image. As a result, we can implicitly obtain training examples beyond what was originally observed in the dataset. Therefore, the deep model is forced to learn subject-independent features, and is generalizable to unseen subjects. SO-Net is end-to-end trainable, and can be flexibly plugged in any CNN model during training. We evaluate the proposed method on two public datasets, BP4D and DISFA. The experiment shows a state-of-the-art performance, demonstrating the effectiveness of the proposed method.
Read more