• Home
  • Search
  • Batch Mode Active Sampling Based on Marginal Probability Distribution Matching
  • Cite Icon75
  • https://doi.org/10.1145/2513092.2513094Copy DOI Icon

Batch Mode Active Sampling Based on Marginal Probability Distribution Matching

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Active Learning is a machine learning and data mining technique that selects the most informative samples for labeling and uses them as training data; it is especially useful when there are large amount of unlabeled data and labeling them is expensive. Recently, batch-mode active learning, where a set of samples are selected concurrently for labeling, based on their collective merit, has attracted a lot of attention. The objective of batch-mode active learning is to select a set of informative samples so that a classifier learned on these samples has good generalization performance on the unlabeled data. Most of the existing batch-mode active learning methodologies try to achieve this by selecting samples based on certain criteria. In this article we propose a novel criterion which achieves good generalization performance of a classifier by specifically selecting a set of query samples that minimize the difference in distribution between the labeled and the unlabeled data, after annotation. We explicitly measure this difference based on all candidate subsets of the unlabeled data and select the best subset. The proposed objective is an NP-hard integer programming optimization problem. We provide two optimization techniques to solve this problem. In the first one, the problem is transformed into a convex quadratic programming problem and in the second method the problem is transformed into a linear programming problem. Our empirical studies using publicly available UCI datasets and two biomedical image databases demonstrate the effectiveness of the proposed approach in comparison with the state-of-the-art batch-mode active learning methods. We also present two extensions of the proposed approach, which incorporate uncertainty of the predicted labels of the unlabeled data and transfer learning in the proposed formulation. In addition, we present a joint optimization framework for performing both transfer and active learning simultaneously unlike the existing approaches of learning in two separate stages, that is, typically, transfer learning followed by active learning. We specifically minimize a common objective of reducing distribution difference between the domain adapted source, the queried and labeled samples and the rest of the unlabeled target domain data. Our empirical studies on two biomedical image databases and on a publicly available 20 Newsgroups dataset show that incorporation of uncertainty information and transfer learning further improves the performance of the proposed active learning based classifier. Our empirical studies also show that the proposed transfer-active method based on the joint optimization framework performs significantly better than a framework which implements transfer and active learning in two separate stages.

Similar Papers
  • Research Article
  • Citations169

Semisupervised SVM batch mode active learning with applications to image retrieval

  • May 01, 2009
  • ACM Transactions on Information Systems
  • Steven C H Hoi +3
  • Research Article
  • Citations9

Batch Mode Active Learning for Node Classification in Assortative and Disassortative Networks

  • Jan 01, 2018
  • IEEE Access
  • Shuqiu Ping +5
  • Research Article
  • Citations350

A comparison review of transfer learning and self-supervised learning: Definitions, applications, advantages and limitations

  • Dec 02, 2023
  • Expert Systems with Applications
  • Zehui Zhao +4
  • Conference Article
  • Citations6

Design and analysis of the WCCI 2010 active learning challenge

  • Jul 01, 2010
  • Isabelle Guyon +3
  • Research Article
  • Citations1

Human Guided Linear Regression With Feature-Level Constraints

  • Apr 29, 2018
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Aubrey Gress +1
  • Conference Article
  • Citations6

Translating Intrusion Alerts to Cyberattack Stages using Pseudo-Active Transfer Learning (PATRL)

  • Oct 04, 2021
  • Stephen Moskal +1
  • Book Chapter
  • Citations1

Transfer Learning with Active Queries for Relational Data Modeling Across Multiple Information Networks

  • Jan 01, 2018
  • Ke-Jia Chen +3
  • Research Article
  • Citations106

Accurate prediction of glaucoma from colour fundus images with a convolutional neural network that relies on active and transfer learning.

  • Jul 25, 2019
  • Acta Ophthalmologica
  • Ruben Hemelings +10
  • Conference Article
  • Citations4

Distributed Active Learning for Image Recognition

  • Mar 01, 2018
  • Shayok Chakraborty
  • Conference Article
  • Citations4

Edge AI for Industry 4.0: An Internet of Things Approach

  • Nov 20, 2020
  • Athanasios Tziouvaras +1
  • Research Article

Exploiting unlabeled data for battery state-of-health estimation using transformer-LSTM neural network with semi-supervised learning

  • Feb 01, 2026
  • Future Batteries
  • Yue Dong +2
  • Research Article
  • Citations60

Reducing Negative Transfer Learning via Clustering for Dynamic Multiobjective Optimization

  • Oct 01, 2022
  • IEEE Transactions on Evolutionary Computation
  • Jianqiang Li +4
  • Conference Article
  • Citations8

Active privileged learning of human activities from weakly labeled samples

  • Sep 01, 2016
  • Michalis Vrigkas +2
  • Book Chapter
  • Citations6

Unsupervised Selective Transfer Learning for Object Recognition

  • Jan 01, 2011
  • Wei-Shi Zheng +2
  • Research Article
  • Citations22

Transfer Learning from Unlabeled Data via Neural Networks

  • Jun 05, 2012
  • Neural Processing Letters
  • Huaxiang Zhang +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.