- Research Article
8
- 10.1016/j.imavis.2022.104444
Dynamic sample weighting for weakly supervised object detection
- Mar 31, 2022
- Image and Vision Computing
- Xuewei Li + 10 more +10
Dynamic sample weighting for weakly supervised object detection
Weakly supervised object detection (WSOD) aims to tackle the object detection problem using only labeled image categories as supervision. A common approach used in WSOD to deal with the lack of localization information is Multiple Instance Learning, and in recent years methods started adopting Multiple Instance Detection Networks (MIDN), which allows training in an end-to-end fashion. In general, these methods work by selecting the best instance from a pool of candidates and then aggregating other instances based on similarity. In this work, we claim that carefully selecting the aggregation criteria can considerably improve the accuracy of the learned detector. We start by proposing an additional refinement step to an existing approach (OICR), which we call refinement knowledge distillation. Then, we present an adaptive supervision aggregation function that dynamically changes the aggregation criteria for selecting boxes related to one of the ground-truth classes, background, or even ignored during the generation of each refinement module supervision. Experiments in Pascal VOC 2007 demonstrate that our Knowledge Distillation and smooth aggregation function significantly improves the performance of OICR in the weakly supervised object detection and weakly supervised object localization tasks. These improvements make the Boosted-OICR competitive again versus other state-of-the-art approaches.
Dynamic sample weighting for weakly supervised object detection
Dynamic sample weighting for weakly supervised object detection
Multiple Instance Complementary Detection and Difficulty Evaluation for Weakly Supervised Object Detection in Remote Sensing Images
Weakly supervised object detection (WSOD) in remote sensing images (RSIs) has attracted lots of attention because it solely employs image-level labels to drive the model training. Most of the WSOD methods incline to mine salient object as positive instance, and the less salient objects are considered as negative instances, which will cause the problem of missing instances. In addition, the quantity of hard and easy instances is usually imbalanced, consequently, the cumulative loss of a large amount of easy instances dominates the training loss, which limits the upper bound of WSOD performance. To handle the first problem, a complementary detection network (CDN) is proposed, which consists of a complementary multiple instance detection network (CMIDN) and a complementary feature learning (CFL) module. The CDN can capture robust complementary information from two basic multiple instance detection networks (MIDNs) and mine more object instances. To handle the second problem, an instance difficulty evaluation metric named instance difficulty score (IDS) is proposed, which is employed as the weight of each instance in the training loss. Consequently, the hard instances will be assigned larger weights according to the IDS, which can improve the upper bound of WSOD performance. The ablation experiments demonstrate that our method significantly increases the baseline method by large margins, i.e. 23.6% (10.2%) mAP and 32.4% (13.1%) CorLoc gains on the NWPU VHR-10.v2 (DIOR) dataset. Our method obtains 58.1% (26.7%) mAP and 72.4% (47.9%) CorLoc on the NWPU VHR-10.v2 (DIOR) dataset, which achieves better performance compared with seven advanced WSOD methods.
Read moreSIOD: Single Instance Annotated Per Category Per Image for Object Detection
Object detection under imperfect data receives great attention recently. Weakly supervised object detection (WSOD) suffers from severe localization issues due to the lack of instance-level annotation, while semi-supervised object detection (SSOD) remains challenging led by the inter-image discrepancy between labeled and unlabeled data. In this study, we propose the Single Instance annotated Object Detection (SIOD), requiring only one instance annotation for each existing category in an image. Degraded from inter-task (WSOD) or inter-image (SSOD) discrepancies to the intra-image discrepancy, SIOD provides more reliable and rich prior knowledge for mining the rest of unlabeled instances and trades off the annotation cost and performance. Under the SIOD setting, we propose a simple yet effective framework, termed Dual-Mining (DMiner), which consists of a Similarity-based Pseudo Label Generating module (SPLG) and a Pixel-level Group Contrastive Learning module (PGCL). SPLG firstly mines latent instances from feature representation space to alleviate the annotation missing problem. To avoid being misled by inaccurate pseudo labels, we propose PGCL to boost the tolerance to false pseudo labels. Extensive experiments on MS COCO verify the feasibility of the SIOD setting and the superiority of the proposed method, which obtains consistent and significant improvements compared to baseline methods and achieves comparable results with fully supervised object detection (FSOD) methods with only 40% instances annotated. Code is available at https://github.com/solicucu/SIOD.
Read moreLearning Object Detectors With Semi-Annotated Weak Labels
For alleviating the human labor associated with annotating the training data for learning object detectors, recent research has focused on semi-supervised object detection (SSOD) and weakly supervised object detection (WSOD) approaches. In SSOD, instead of annotating all the instances in the whole training set, people only need to annotate the part of the training instances using bounding boxes. In WSOD, people need to annotate the image-level tags on all training images to indicate the object categories contained by the corresponding images since more detailed bounding box annotations are no longer needed. Along this line of research, this paper makes a further step to alleviate the human labor in annotating training data, leading to the problem of object detection with semi-annotated weak labels (ODSAWLs). Instead of labeling image-level tags on all training images, ODSAWL only needs the image-level tags for a small portion of the training images, and then, the object detectors can be learned from a small portion of the weakly-labeled training images and from the remaining unlabeled training images. To address such a challenging problem, this paper proposes a cross model co-training framework that collaborates an object localizer and a tag generator in an alternative optimization procedure. Specifically, during the learning procedure, these two (deep) models can transfer the needed knowledge (including labels and visual patterns) between each other. The whole learning procedure is accomplished in a few stages under the guidance of a progressive learning curriculum. To demonstrate the effectiveness of the proposed approach, we implement the comprehensive experiments on three benchmark datasets, where the obtained experimental results are quite encouraging. Notably, by using only about 15% weakly labeled training images, the proposed approach can effectively approach, or even outperform, the state-of-the-art WSOD methods.
Read moreFew-shot Weakly-Supervised Object Detection via Directional Statistics
Detecting novel objects from few examples has become an emerging topic in computer vision recently. However, current methods need fully annotated training images to learn new object categories which limits their applicability in real world scenarios such as field robotics. In this work, we propose a probabilistic multiple-instance learning approach for few-shot Common Object Localization (COL) and few-shot Weakly Supervised Object Detection (WSOD). In these tasks, only image-level labels, which are much cheaper to acquire, are available. We find that operating on features extracted from the last layer of a pretrained Faster-RCNN is more effective compared to previous episodic learning based few-shot COL methods. Our model simultaneously learns the distribution of the novel objects and localizes them via expectation-maximization steps. As a probabilistic model, we employ von Mises-Fisher (vMF) distribution which captures the semantic information better than Gaussian distribution when applied to the pre-trained embedding space. When the novel objects are localized, we utilize them to learn a linear appearance model to detect novel classes in new images. Our extensive experiments show that the proposed method, despite being simple, outperforms strong baselines in few-shot COL and WSOD, as well as large-scale WSOD tasks.
Read moreDeep Learning for Weakly-Supervised Object Detection and Localization: A Survey
Deep Learning for Weakly-Supervised Object Detection and Localization: A Survey
Incorporating the Completeness and Difficulty of Proposals Into Weakly Supervised Object Detection in Remote Sensing Images
Weakly supervised object detection (WSOD) in remote sensing images (RSI) only require image-level labels to detect various objects. Most of the WSOD methods incline to capture the most discriminative parts of object rather than the entire object, and the number of easy and hard samples is imbalanced. To address the first problem, a novel metric named objectness score (OS) is proposed and incorporated into the training loss of our WSOD model. The OS is consisted of the traditional class confidence score (CCS) and the object completeness prior score (OCPS). The CCS can provide the probability that a proposal belongs to a certain class, and the OCPS can quantify the completeness that a proposal covers the entire object. Therefore, the samples which cover the entire object with high class confidences will be assigned large weight in the training loss through OS. To handle the second problem, a novel metric named difficulty evaluation score (DES) is proposed and also incorporated into the training loss. The DES is calculated by using the entropy of confidence score vector of each proposal and is used to quantify how difficult a proposal can be identified correctly, consequently, the hard samples will also be assigned large weight in the training loss through DES. The ablation experiments on two RSI datasets verify the effectiveness of the proposed OS and DES. The comprehensive quantitative and subjective evaluations demonstrate that our method inclines to detect the entire object accurately, and surpasses seven state-of-the-art WSOD methods.
Read moreDynamic Informative Proposal-Based Iterative Training Procedure for Weakly Supervised Object Detection in Remote Sensing Images
Weakly supervised object detection (WSOD) is an increasingly important task in remote sensing images. However, mainstream WSOD methods often rely on low-quality proposals due to the complex backgrounds of remote sensing images. Moreover, applying strong data augmentations directly in WSOD methods can introduce significant noise, which can hinder training procedures that rely only on image-level ground truth labels. To address these issues, we propose a dynamic informative proposal-based iterative WSOD training procedure. Specifically, we implement an Informative Proposal Reconstruction (IPR) method to generate more informative proposals dynamically. We also use a Proposal-Based Contrastive Learning (PBCL) technique to steadily improve the quality of generated proposals. In addition, we employ a Pseudo Label Learning-based Multi-Stage training procedure (PLMS) to progressively improve the quality of new informative proposals while alleviating the noise catastrophe caused by strong data augmentations. Extensive experiments demonstrate the effectiveness of our proposed method in generating higher quality proposals and enhancing model generalization. Our method achieves state-of-the-art results in optical remote sensing images (DIOR), Northwestern Polytechnical University (NWPU) VHR-10.v2, and HRSC2016 datasets.
Read moreEnabling Deep Residual Networks for Weakly Supervised Object Detection
Weakly supervised object detection (WSOD) has attracted extensive research attention due to its great flexibility of exploiting large-scale image-level annotation for detector training. Whilst deep residual networks such as ResNet and DenseNet have become the standard backbones for many computer vision tasks, the cutting-edge WSOD methods still rely on plain networks, e.g., VGG, as backbones. It is indeed not trivial to employ deep residual networks for WSOD, which even shows significant deterioration of detection accuracy and non-convergence. In this paper, we discover the intrinsic root with sophisticated analysis and propose a sequence of design principles to take full advantages of deep residual learning for WSOD from the perspectives of adding redundancy, improving robustness and aligning features. First, a redundant adaptation neck is key for effective object instance localization and discriminative feature learning. Second, small-kernel convolutions and MaxPool down-samplings help improve the robustness of information flow, which gives finer object boundaries and make the detector more sensitivity to small objects. Third, dilated convolution is essential to align the proposal features and exploit diverse local information by extracting high-resolution feature maps. Extensive experiments show that the proposed principles enable deep residual networks to establishes new state-of-the-arts on PASCAL VOC and MS COCO.
Read moreObject Instance Mining for Weakly Supervised Object Detection
Weakly supervised object detection (WSOD) using only image-level annotations has attracted growing attention over the past few years. Existing approaches using multiple instance learning easily fall into local optima, because such mechanism tends to learn from the most discriminative object in an image for each category. Therefore, these methods suffer from missing object instances which degrade the performance of WSOD. To address this problem, this paper introduces an end-to-end object instance mining (OIM) framework for weakly supervised object detection. OIM attempts to detect all possible object instances existing in each image by introducing information propagation on the spatial and appearance graphs, without any additional annotations. During the iterative learning process, the less discriminative object instances from the same class can be gradually detected and utilized for training. In addition, we design an object instance reweighted loss to learn larger portion of each object instance to further improve the performance. The experimental results on two publicly available databases, VOC 2007 and 2012, demonstrate the efficacy of proposed approach.
Read moreSelf-Training-Based Semantic-Balanced Network for Weakly Supervised Object Detection in Remote-Sensing Images
A weakly supervised object detection (WSOD) task is to train a detector with only image-level labels provided. Except for the training difficulty introduced by weaker annotations, the inherent complexity of the remote-sensing images (RSIs) also adds to the challenge. To boost the detector’s localization accuracy, we aim to exploit more semantic information contained in images and help improve the general robustness of the model. Noticing previous methods tend to focus on the most discriminative part of an object, we design a self-training-based network that leverages local semantic features. To this end, we develop a semantic-balanced localization module (SBLM) that distinguishes foreground from background and accurate proposals from incomplete ones, by leveraging a balance of region of interest (ROI) and its context information. Moreover, we find that the self-training strategy highly relies on the quality of pseudo-ground-truth boxes. Motivated by this possible lack of robustness, we design a comprehensive clustering module (CCM) and saliency-based proposal filtering (SPF) module that select pseudo-ground truth more comprehensively under supervision. To be more specific, CCM aims to reduce the arbitrariness during assigning pseudo-labels by considering multiple categorical vectors simultaneously. Salient object detection (SOD) is applied in the SPF module to help evaluate the quality of the chosen pseudo-ground-truth boxes. The detection performance is significantly boosted with the proposed method. Extensive experiments conducted on the NWPU VHR-10.v2 dataset and the DIOR dataset validate that the proposed model outperforms the previous state-of-the-art methods favorably with an mAP of 64.9% and 28.1%, respectively.
Read moreSAM-Induced Pseudo Fully Supervised Learning for Weakly Supervised Object Detection in Remote Sensing Images
Weakly supervised object detection (WSOD) in remote sensing images (RSIs) aims to detect high-value targets by solely utilizing image-level category labels; however, two problems have not been well addressed by existing methods. Firstly, the seed instances (SIs) are mined solely relying on the category score (CS) of each proposal, which is inclined to concentrate on the most salient parts of the object; furthermore, they are unreliable because the robustness of the CS is not sufficient due to the fact that the inter-category similarity and intra-category diversity are more serious in RSIs. Secondly, the localization accuracy is limited by the proposals generated by the selective search or edge box algorithm. To address the first problem, a segment anything model (SAM)-induced seed instance-mining (SSIM) module is proposed, which mines the SIs according to the object quality score, which indicates the comprehensive characteristic of the category and the completeness of the object. To handle the second problem, a SAM-based pseudo-ground truth-mining (SPGTM) module is proposed to mine the pseudo-ground truth (PGT) instances, for which the localization is more accurate than traditional proposals by fully making use of the advantages of SAM, and the object-detection heads are trained by the PGT instances in a fully supervised manner. The ablation studies show the effectiveness of the SSIM and SPGTM modules. Comprehensive comparisons with 15 WSOD methods demonstrate the superiority of our method on two RSI datasets.
Read moreSelf-supervised object detection from audio-visual correspondence
We tackle the problem of learning object detectors without supervision. Differently from weakly-supervised object detection, we do not assume image-level class labels. Instead, we extract a supervisory signal from audio-visual data, using the audio component to "teach" the object detector. While this problem is related to sound source localisation, it is considerably harder because the detector must classify the objects by type, enumerate each instance of the object, and do so even when the object is silent. We tackle this problem by first designing a self-supervised framework with a contrastive objective that jointly learns to classify and localise objects. Then, without using any supervision, we simply use these self-supervised labels and boxes to train an image-based object detector. With this, we outperform previous unsupervised and weakly-supervised detectors for the task of object detection and sound source localization. We also show that we can align this detector to ground-truth classes with as little as one label per pseudo-class, and show how our method can learn to detect generic objects that go beyond instruments, such as airplanes and cats.
Read morePistonNet: Object Separating From Background by Attention for Weakly Supervised Ship Detection
Object detection under weakly supervised learning is a challenging issue. In remote sensing ship detection task, the weakly supervised learning method requires that the training set only has image-level class annotations. In the absence of location information, it is difficult to locate ships and extract features. Moreover, when the detector extracts more than one candidate region, the image-level annotation makes it difficult to determine whether the candidate regions are multiple ships or mixed with the detected background. To address these issues, this article analyzes the interaction between class information and location information, and proposes a weakly supervised detection method, PistonNet, based on Data-efficient image Transformers (DeiT). PistonNet proposes an artificial point, which is inserted into the feature map. Artificial point can suppress the background and enhance the object by interfering in the weight distribution of object and background in Self-Attention calculation. PistonNet also proposes joint confidence probability to improve detection accuracy. Experiments on the GF1-LRSD and NWPU VHR-10 dataset show that the proposed method boosts detection performance effectively. The main improvements of PistonNet as well as the contributions of this article are threefold. Firstly, PistonNet is a specific weakly supervised object detection (WSOD) method for single-class detection, which provides an innovative approach to classifying target and background. Secondly, PistonNet reaches the level of advanced supervised detectors on detection accuracy with fewer parameters. Finally, the objects’ locations detected by PistonNet are obtained by segmenting regions on the heatmap. PistonNet’s background suppression ability makes it free from dependence on segmentation threshold.
Read moreMultiple instance boosting with global smoothness regularization
In multiple instance learning, the training set consists of labeled bags that include unlabeled instances, and the target is to predict the labels of unseen bags. A bag is labeled positive only if it contains at least one positive instance, otherwise it is a negative bag. Over the past years, many popular machine learning algorithms have been adapted to tackle the multiple instance learning problems. In this paper, to train a discriminative multiple instance classifier which generalize well, we present a boosting approach with global smoothness regularization, in which the weak learners are either hyper balls with the center at the instance of positive bags or random projection decision stumps. Experimental results show that our proposed algorithm is comparable to the classical Diverse Density algorithm on some multiple instance learning benchmark datasets.
Read more