- Research Article
334
- 10.1016/j.patrec.2013.06.010
A bagging SVM to learn from positive and unlabeled examples
- Jun 25, 2013
- Pattern Recognition Letters
- F Mordelet + 1 more +1
A bagging SVM to learn from positive and unlabeled examples
This paper studies the problem of Positive Unlabeled learning (PU learning), where positive and unlabeled examples are used for training. Naive Bayes (NB) and Tree Augmented Naive Bayes (TAN) have been extended to PU learning algorithms (PNB and PTAN). However, they require user-specified parameter, which is difficult for the user to provide in practice. We estimate this parameter following [2] by taking the selected completely at random assumption and reformulate these two algorithms with this assumption. Furthermore, based on supervised algorithms Averaged One-Dependence Estimators (AODE), Hidden Naive Bayes (HNB) and Full Bayesian network Classifier (FBC), we extend these algorithms to PU learning algorithms (PAODE, PHNB and PFBC respectively). Experimental results on 20 UCI datasets show that the performance of the Bayesian algorithms for PU learning are comparable to corresponding supervised ones in most cases. Additionally, PNB and PFBC are more robust against unlabeled data, and PFBC generally performs the best.
A bagging SVM to learn from positive and unlabeled examples
A bagging SVM to learn from positive and unlabeled examples
Two birds with one stone: Classifying positive and unlabeled examples on uncertain data streams
Two birds with one stone: Classifying positive and unlabeled examples on uncertain data streams
Absolute Value Inequality SVM for the PU Learning Problem
Positive and unlabeled learning (PU learning) is a significant binary classification task in machine learning; it focuses on training accurate classifiers using positive data and unlabeled data. Most of the works in this area are based on a two-step strategy: the first step is to identify reliable negative examples from unlabeled examples, and the second step is to construct the classifiers based on the positive examples and the identified reliable negative examples using supervised learning methods. However, these methods always underutilize the remaining unlabeled data, which limits the performance of PU learning. Furthermore, many methods require the iterative solution of the formulated quadratic programming problems to obtain the final classifier, resulting in a large computational cost. In this paper, we propose a new method called the absolute value inequality support vector machine, which applies the concept of eccentricity to select reliable negative examples from unlabeled data and then constructs a classifier based on the positive examples, the selected negative examples, and the remaining unlabeled data. In addition, we apply a hyperparameter optimization technique to automatically search and select the optimal parameter values in the proposed algorithm. Numerical experimental results on ten real-world datasets demonstrate that our method is better than the other three benchmark algorithms.
Read morePositive-unlabeled learning in bioinformatics and computational biology: a brief review.
Conventional supervised binary classification algorithms have been widely applied to address significant research questions using biological and biomedical data. This classification scheme requires two fully labeled classes of data (e.g. positive and negative samples) to train a classification model. However, in many bioinformatics applications, labeling data is laborious, and the negative samples might be potentially mislabeled due to the limited sensitivity of the experimental equipment. The positive unlabeled (PU) learning scheme was therefore proposed to enable the classifier to learn directly from limited positive samples and a large number of unlabeled samples (i.e. a mixture of positive or negative samples). To date, several PU learning algorithms have been developed to address various biological questions, such as sequence identification, functional site characterization and interaction prediction. In this paper, we revisit a collection of 29 state-of-the-art PU learning bioinformatic applications to address various biological questions. Various important aspects are extensively discussed, including PU learning methodology, biological application, classifier design and evaluation strategy. We also comment on the existing issues of PU learning and offer our perspectives for the future development of PU learning applications. We anticipate that our work serves as an instrumental guideline for a better understanding of the PU learning framework in bioinformatics and further developing next-generation PU learning frameworks for critical biological applications.
Read moreEstimating classification accuracy in positive-unlabeled learning: characterization and correction strategies
Accurately estimating performance accuracy of machine learning classifiers is of fundamental importance in biomedical research with potentially societal consequences upon the deployment of bestperforming tools in everyday life. Although classification has been extensively studied over the past decades, there remain understudied problems when the training data violate the main statistical assumptions relied upon for accurate learning and model characterization. This particularly holds true in the open world setting where observations of a phenomenon generally guarantee its presence but the absence of such evidence cannot be interpreted as the evidence of its absence. Learning from such data is often referred to as positive-unlabeled learning, a form of semi-supervised learning where all labeled data belong to one (say, positive) class. To improve the best practices in the field, we here study the quality of estimated performance in positive-unlabeled learning in the biomedical domain. We provide evidence that such estimates can be wildly inaccurate, depending on the fraction of positive examples in the unlabeled data and the fraction of negative examples mislabeled as positives in the labeled data. We then present correction methods for four such measures and demonstrate that the knowledge or accurate estimates of class priors in the unlabeled data and noise in the labeled data are sufficient for the recovery of true classification performance. We provide theoretical support as well as empirical evidence for the efficacy of the new performance estimation methods.
Read morePositive Unlabeled Learning for Deceptive Reviews Detection
Deceptive reviews detection has attracted significant attention from both business and research communities. However, due to the difficulty of human labeling needed for supervised learning, the problem remains to be highly challenging. This paper proposed a novel angle to the problem by modeling PU (positive unlabeled) learning. A semi-supervised model, called mixing population and individual property PU learning (MPIPUL), is proposed. Firstly, some reliable negative examples are identified from the unlabeled dataset. Secondly, some representative positive examples and negative examples are generated based on LDA (Latent Dirichlet Allocation). Thirdly, for the remaining unlabeled examples (we call them spy examples), which can not be explicitly identified as positive and negative, two similarity weights are assigned, by which the probability of a spy example belonging to the positive class and the negative class are displayed. Finally, spy examples and their similarity weights are incorporated into SVM (Support Vector Machine) to build an accurate classifier. Experiments on gold-standard dataset demonstrate the effectiveness of MPIPUL which outperforms the state-of-the-art baselines.
Read moreA Tax Evasion Detection Method Based on Positive and Unlabeled Learning with Network Embedding Features
Tax evasion detection has a crucial role in addressing tax revenue loss. In the real world, an accessed tax dataset only contains a small number of labeled taxpayers who evade tax (positive samples) and a large number of unlabeled taxpayers who either evade tax or do not evade tax. It is difficult to address this issue due to this nontraditional dataset. In addition, the basic features of taxpayers designed according to tax experts’ domain knowledge and experience are very limited to determining whether taxpayers evade tax. These limitations motivate the contribution of this work. In this paper, we argue that the tax evasion detection task in the real world should be formalized as a positive unlabeled (PU) learning problem. We propose a novel tax evasion detection method based on PU learning with Network Embedding features (PUNE). PUNE effectively detects tax evasion based on basic features and transaction network features that are extracted by a network embedding algorithm. Moreover, PUNE can work well even under label noise. To evaluate the effectiveness of PUNE, we conduct experimental tests on a real-world tax dataset. The results demonstrate that PUNE can significantly improve the performance of tax evasion detection.
Read moreC-PUGP: A cluster-based positive unlabeled learning method for disease gene prediction and prioritization
C-PUGP: A cluster-based positive unlabeled learning method for disease gene prediction and prioritization
Learning from positive and unlabeled data: a survey
Learning from positive and unlabeled data or PU learning is the setting where a learner only has access to positive examples and unlabeled data. The assumption is that the unlabeled data can contain both positive and negative examples. This setting has attracted increasing interest within the machine learning literature as this type of data naturally arises in applications such as medical diagnosis and knowledge base completion. This article provides a survey of the current state of the art in PU learning. It proposes seven key research questions that commonly arise in this field and provides a broad overview of how the field has tried to address them.
Read morePositive and unlabeled learning from hospital administrative data: a novel approach to identify sepsis cases.
In positive and unlabeled (PU) learning problems, only positive examples are labeled. Unlabeled data contain both positive and negative examples. Studies show that positive examples of (secondary) diagnoses, and clinical conditions, such as sepsis, are present in unlabeled hospital administrative data, potentially distorting hospital reimbursement systems, and negatively affecting hospitals' revenue and profitability. We investigate whether PU learning is suitable for improving the quality of hospital administrative data. We train three models on 313,434 hospital cases using hospital cost features: two based on the two-step "spy" approach and one using a robust PU learning method. For model evaluation, we rely exclusively on positive examples due to the PU setting. To further assess model performance, we perform an external validity check: We relabel unlabeled sepsis cases, derive new sepsis rates, and compare them to those reported in medical record review studies. All models identify true positives well in unseen data. External validity checks show, however, that only the robust PU learner effectively discriminates between positives and negatives in the unlabeled data, yielding new sepsis rates within the range of sepsis rates reported in medical record review studies. PU learning can improve the quality of hospital administrative data, but its effectiveness depends strongly on the choice of learning approach and classifier. The output of a PU learner can potentially improve hospital reimbursement systems, hospital revenue and profitability management, and sensitivity analyses in healthcare management science, health economics, health services research, and disease surveillance.
Read moreLearning classifiers without negative examples: A reduction approach
The problem of PU Learning, i.e., learning classifiers with positive and unlabelled examples (but not negative examples), is very important in information retrieval and data mining. We address this problem through a novel approach: reducing it to the problem of learning classifiers for some meaningful multivariate performance measures. In particular, we show how a powerful machine learning algorithm, support vector machine, can be adapted to solve this problem. The effectiveness and efficiency of the proposed approach have been confirmed by our experiments on three real-world datasets.
Read moreLearning Bayesian network classifiers for facial expression recognition both labeled and unlabeled data
Understanding human emotions is one of the necessary skills for the computer to interact intelligently with human users. The most expressive way humans display emotions is through facial expressions. In this paper, we report on several advances we have made in building a system for classification of facial expressions from continuous video input. We use Bayesian network classifiers for classifying expressions from video. One of the motivating factor in using the Bayesian network classifiers is their ability to handle missing data, both during inference and training. In particular, we are interested in the problem of learning with both labeled and unlabeled data. We show that when using unlabeled data to learn classifiers, using correct modeling assumptions is critical for achieving improved classification performance. Motivated by this, we introduce a classification driven stochastic structure search algorithm for learning the structure of Bayesian network classifiers. We show that with moderate size labeled training sets and large amount of unlabeled data, our method can utilize unlabeled data to improve classification performance. We also provide results using the Naive Bayes (NB) and the Tree-Augmented Naive Bayes (TAN) classifiers, showing that the two can achieve good performance with labeled training sets, but perform poorly when unlabeled data are added to the training set.
Read morePositive-unlabeled learning for open set domain adaptation
Positive-unlabeled learning for open set domain adaptation
Learning Classifiers on Positive and Unlabeled Data with Policy Gradient
Existing algorithms aiming to learn a binary classifier from positive (P) and unlabeled (U) data generally require estimating the class prior or label noises ahead of building a classification model. However, the estimation and classifier learning are normally conducted in a pipeline instead of being jointly optimized. In this paper, we propose to alternatively train the two steps using reinforcement learning. Our proposal adopts a policy network to adaptively make assumptions on the labels of unlabeled data, while a classifier is built upon the output of the policy network and provides rewards to learn a better strategy. The dynamic and interactive training between the policy maker and the classifier can exploit the unlabeled data in a more effective manner and yield a significant improvement on the classification performance. Furthermore, we present two different approaches to represent the actions sampled from the policy. The first approach considers continuous actions as soft labels, while the other uses discrete actions as hard assignment of labels for unlabeled examples. We validate the effectiveness of the proposed method on two benchmark datasets as well as one e-commerce dataset. The result shows the proposed method is able to consistently outperform state-of-the-art methods in various settings.
Read moreA Novel Approach for Cross-Selling Insurance Products Using Positive Unlabelled Learning
Successful cross-selling of products is a key goal of companies operating within the insurance industry. Choosing the right customer to approach for cross-purchase opportunities has a direct effect on both decreasing customer churn rate and increasing revenue. Unlike sales data of general products, insurance sales data typically contains only a few products (e.g., private medical insurance, life insurance, etc), it is highly imbalanced with a vast majority of customers with no cross-purchasing information, highly noisy due to varying purchase behaviour between different customers, and has no ground truth for knowing if the majority customers are truly non-cross-sell customers or they are missed opportunities. These data challenges render the building of machine learning models for accurately identifying potential cross-sell customers extremely difficult. This paper proposes a novel approach to solve this challenging problem of cross-sell customer identification using Positive Unlabelled (PU) learning in conjunction with advanced feature engineering on customer demographic data and unstructured customer question-response texts through topic modelling. We implement a bagging approach to iteratively learn the positive samples (the confirmed cross-sells) alongside random sub-samples of the unlabelled set. The proposed approach is extensively evaluated on real insurance data that has been newly collected from a leading insurance company for this study. Evaluation results demonstrate that our approach can successfully identify new potential opportunities for likely cross-sell customers.
Read more