- Research Article
119
- 10.1016/j.neucom.2020.04.040
Revisiting metric learning for few-shot image classification
- Apr 21, 2020
- Neurocomputing
- Xiaomeng Li + 4 more +4
Revisiting metric learning for few-shot image classification
Few-shot learning aims to learn a classifier to recognize unseen classes with limited labeled examples. The scarcity of training data remains a challenging problem in few-shot classification. Many few-shot learning methods focus on the structure of meta-network models and there is often no dedicated study of meta-knowledge representation. In this paper, a meta-learning strategy is introduced and a meta-network NAM Net is proposed. Taking the advantage of 'learning to learn', the model acquires meta-knowledge via few-shot classification tasks and applies it to new few-shot scenarios. Specifically, the meta-learner learns the best parameters of the feature extraction and similarity metric modules. The distance between the support feature and query feature is obtained by a learnable metric function, which leads to the classification result. The base learner migrates the meta-knowledge to the target class to perform classification in a new few-shot episode. Moreover, the Normalization-based Attention Module is adapted to feature extractor to enhance meta-knowledge representation. Compared with few-shot learning benchmarks, NAM Net is effective and achieves higher accuracy in both 5-way 1-shot and 5-shot classification tasks.
Revisiting metric learning for few-shot image classification
Revisiting metric learning for few-shot image classification
Coarse-to-fine pseudo supervision guided meta-task optimization for few-shot object classification
Coarse-to-fine pseudo supervision guided meta-task optimization for few-shot object classification
ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot Learning
Self-supervised learning (SSL) techniques have recently been integrated into the few-shot learning (FSL) framework and have shown promising results in improving the few-shot image classification performance. However, existing SSL approaches used in FSL typically seek the supervision signals from the global embedding of every single image. Therefore, during the episodic training of FSL, these methods cannot capture and fully utilize the local visual information in image samples and the data structure information of the whole episode, which are beneficial to FSL. To this end, we propose to augment the few-shot learning objective with a novel self-supervised Episodic Spatial Pretext Task (ESPT). Specifically, for each few-shot episode, we generate its corresponding transformed episode by applying a random geometric transformation to all the images in it. Based on these, our ESPT objective is defined as maximizing the local spatial relationship consistency between the original episode and the transformed one. With this definition, the ESPT-augmented FSL objective promotes learning more transferable feature representations that capture the local spatial features of different images and their inter-relational structural information in each input episode, thus enabling the model to generalize better to new categories with only a few samples. Extensive experiments indicate that our ESPT method achieves new state-of-the-art performance for few-shot image classification on three mainstay benchmark datasets. The source code will be available at: https://github.com/Whut-YiRong/ESPT.
Read moreFrom patch, sample to domain: Capture geometric structures for few-shot learning
From patch, sample to domain: Capture geometric structures for few-shot learning
Multi-Learner Based Deep Meta-Learning for Few-Shot Medical Image Classification.
Few-shot learning (FSL) is promising in the field of medical image analysis due to high cost of establishing high-quality medical datasets. Many FSL approaches have been proposed in natural image scenes. However, present FSL methods are rarely evaluated on medical images and the FSL technology applicable to medical scenarios need to be further developed. Meta-learning has supplied an optional framework to address the challenging FSL setting. In this paper, we propose a novel multi-learner based FSL method for multiple medical image classification tasks, combining meta-learning with transfer-learning and metric-learning. Our designed model is composed of three learners, including auto-encoder, metric-learner and task-learner. In transfer-learning, all the learners are trained on the base classes. In the ensuing meta-learning, we leverage multiple novel tasks to fine-tune the metric-learner and task-learner in order to fast adapt to unseen tasks. Moreover, to further boost the learning efficiency of our model, we devised real-time data augmentation and dynamic Gaussian disturbance soft label (GDSL) scheme as effective generalization strategies of few-shot classification tasks. We have conducted experiments for three-class few-shot classification tasks on three newly-built challenging medical benchmarks, BLOOD, PATH and CHEST. Extensive comparisons to related works validated that our method achieved top performance both on homogeneous medical datasets and cross-domain datasets.
Read moreA Closer Look at Few-Shot Classification with Many Novel Classes
Few-shot learning (FSL) is designed to equip models with the capability to quickly adapt to new, unseen domains in open-world scenarios. However, there is a notable discrepancy between the multitude of new concepts encountered in the open world and the limited scale of existing FSL studies, which focus predominantly on a small number of novel classes. This limitation hinders the practical implementation of FSL in real-world situations. To address this issue, we introduce a novel problem called Few-Shot Learning with Many Novel Classes (FSL-MNC), which expands the number of novel classes more than 500 times compared to traditional FSL settings. This new challenge presents two main difficulties: increased computational load during meta-training and reduced classification accuracy due to the larger number of classes during meta-testing. To tackle these problems, we introduce the Simple Hierarchy Pipeline (SHA-Pipeline). In response to the inefficiency of traditional Episode Meta-Learning (EML) protocols, we redesign a more efficient meta-training strategy to manage the increased number of novel classes. Moreover, to distinguish distinct semantic features across a broad array of novel classes, we effectively reconstruct and utilize class hierarchy information during meta-testing. Our experiments demonstrate that the SHA-Pipeline substantially outperforms both the ProtoNet baseline and current leading alternatives across various numbers of novel classes.
Read moreHow to make use of pretrained models in few-shot classification
Few-shot learning(FSL) aims to generalize model to novel categoeries by few labelled samples, which is challenging for machine. Large-scaled pretrained models, especially vision transformers achieve excellent performances benefiting from numerous and diverse data. Researchers have exploited pretrained models in few-shot classification by simply updating the whole parameters and finetuning on few samples. In this paper, we explore two methods: vision prompt tuning and a reparameterization method called ‘scaling&&shift’ to leverage pretrained models in few-shot classification. Vision prompt tuning is for vision transformer only and we first evaluate the method in few-shot setting. ‘Scaling&&shift’ is originally applied in convolution neural networks(CNN). We extend it to vision transformer. The two methods are evaluated on standard benchmarks such as miniImageNet, CUB, CIFAR-FS, clipart and sketch. The results show that ‘scaling&&shift’ reaches the same level compared to updating the whole parameters. Vision prompt tuning is 0%~5% lower than updating the whole parameters over five datasets while it has quite smaller amount of parameters updated.
Read moreTST_MFL: Two-stage training based metric fusion learning for few-shot image classification
TST_MFL: Two-stage training based metric fusion learning for few-shot image classification
ConASD: Contrastive Few Shot Learning for Detecting Autism Spectrum Disorder via Eye Tracking Scanpath
Detecting Autism Spectrum Disorder (ASD) using Eye Tracking (ET) datasets is a challenging task and has been a long-standing problem. Recently, there has been a trend of developing ASD diagnosis models based on machine learning (ML), especially deep learning techniques. In this paper, we show that these existing methods still struggle to make accurate diagnoses in few-shot learning (FSL) settings, where the data available for training is limited in amount and imbalanced in nature. To address this challenge, we propose a model, named ConASD, for effective diagnosis of ASD under the FSL setting. The proposed model is a two-stage framework: it first trains an encoder for ET images using supervised contrastive learning, followed by fine-tuning a classifier for final diagnosis. With the contrastive learning strategy, the pre-trained encoder can better capture the discriminative features of the eye-tracking images, even with limited training data, and ultimately leads to better diagnosis accuracy and better generalization to unseen data. We evaluate the proposed ConASD model using two real-world ET datasets. The results demonstrate that ConASD outperforms existing approaches, particularly in few-shot scenarios, by up-to 7% improvement in terms of F1 scores. The results in this paper highlight the potential of using contrastive learning as a powerful tool, particularly in real-world medical scenarios where class imbalance is frequent and the data is limited.
Read moreContrastive Representation for Dermoscopy Image Few-Shot Classification
In the field of few-shot learning, different methods are proposed to optimize the model by changing the network structure or optimizing the algorithm. Although, by designing the end-to-end algorithms, the classification performance on a specific task can be improved. But when the amount of data is limited, it is still difficult to obtain generalization ability for different tasks. At the same time, a single loss function in the end-to-end training is always ineffective, because it is difficult for a shallow network to learn an effective image feature representation from complex natural images. In this paper, a constructive representation algorithm for few-shot learning without end-to-end training is proposed, which is suitable for few-shot learning in the natural images classification task. Through self-supervised representation learning, the proposed encoding model generates an effective feature representation. Then, a few-shot learning model is further constructed and trained for supervised classification tasks. Without the end-to-end training, the proposed learning method at different training stages use different loss functions. In the experiments, on the classification task of the public competition data set ISIC2018[11], our method has a 30% performance improvement over the state-of-the-art methods.
Read moreGraph-Based Domain Adaptation Few-Shot Learning for Hyperspectral Image Classification
Due to a lack of labeled samples, deep learning methods generally tend to have poor classification performance in practical applications. Few-shot learning (FSL), as an emerging learning paradigm, has been widely utilized in hyperspectral image (HSI) classification with limited labeled samples. However, the existing FSL methods generally ignore the domain shift problem in cross-domain scenes and rarely explore the associations between samples in the source and target domain. To tackle the above issues, a graph-based domain adaptation FSL (GDAFSL) method is proposed for HSI classification with limited training samples, which utilizes the graph method to guide the domain adaptation learning process in a uniformed framework. First, a novel deep residual hybrid attention network (DRHAN) is designed to extract discriminative embedded features efficiently for few-shot HSI classification. Then, a graph-based domain adaptation network (GDAN), which combines graph construction with domain adversarial strategy, is proposed to fully explore the domain correlation between source and target embedded features. By utilizing the fully explored domain correlations to guide the domain adaptation process, a domain invariant feature metric space is learned for few-shot HSI classification. Comprehensive experimental results conducted on three public HSI datasets demonstrate that GDAFSL is superior to the state-of-the-art with a small sample size.
Read moreGenerally Boosting Few-Shot Learning with HandCrafted Features
Existing Few-Shot Learning (FSL) methods predominantly focus on developing different types of sophisticated models to extract the transferable prior knowledge for recognizing novel classes, while they almost pay less attention to the feature learning part in FSL which often simply leverage some well-known CNN as the feature learner. However, feature is the core medium for encoding such transferable knowledge. Feature learning is easy to be trapped in the over-fitting particularly in the scarcity of the training data, and thereby degenerates the performances of FSL. The handcrafted features, such as Histogram of Oriented Gradient (HOG) and Local Binary Pattern (LBP), have no requirement on the amount of training data, and used to perform quite well in many small-scale data scenarios, since their extractions involve no learning process, and are mainly based on the empirically observed and summarized prior feature engineering knowledge. In this paper, we intend to develop a general and simple approach for generally boosting FSL via exploiting such prior knowledge in the feature learning phase. To this end, we introduce two novel handcrafted feature regression modules, namely HOG and LBP regression, to the feature learning parts of deep learning-based FSL models. These two modules are separately plugged into the different convolutional layers of backbone based on the characteristics of the corresponding handcrafted features to guide the backbone optimization from different feature granularity, and also ensure that the learned feature can encode the handcrafted feature knowledge which improves the generalization ability of feature and alleviate the over-fitting of the models. Three recent state-of-the-art FSL approaches are leveraged for examining the effectiveness of our method. Extensive experiments on miniImageNet, CIFAR-FS and FC100 datasets show that the performances of all these FSL approaches are well boosted via applying our method on all three datasets. Our codes and models have been released.
Read moreDynamic Subcluster-Aware Network for Few-Shot Skin Disease Classification.
This article addresses the problem of few-shot skin disease classification by introducing a novel approach called the subcluster-aware network (SCAN) that enhances accuracy in diagnosing rare skin diseases. The key insight motivating the design of SCAN is the observation that skin disease images within a class often exhibit multiple subclusters, characterized by distinct variations in appearance. To improve the performance of few-shot learning (FSL), we focus on learning a high-quality feature encoder that captures the unique subclustered representations within each disease class, enabling better characterization of feature distributions. Specifically, SCAN follows a dual-branch framework, where the first branch learns classwise features to distinguish different skin diseases, and the second branch aims to learn features, which can effectively partition each class into several groups so as to preserve the subclustered structure within each class. To achieve the objective of the second branch, we present a cluster loss to learn image similarities via unsupervised clustering. To ensure that the samples in each subcluster are from the same class, we further design a purity loss to refine the unsupervised clustering results. We evaluate the proposed approach on two public datasets for few-shot skin disease classification. The experimental results validate that our framework outperforms the state-of-the-art methods by around 2%-5% in terms of sensitivity, specificity, accuracy, and F1-score on the SD-198 and Derm7pt datasets.
Read moreMulti-granularity Recurrent Attention Graph Neural Network for Few-Shot Learning
Few-shot learning aims to learn a classifier that classifies unseen classes well with limited labeled samples. Existing meta learning-based works, whether graph neural network or other baseline approaches in few-shot learning, has benefited from the meta-learning process with episodic tasks to enhance the generalization ability. However, the performance of meta-learning is greatly affected by the initial embedding network, due to the limited number of samples. In this paper, we propose a novel Multi-granularity Recurrent Attention Graph Neural Network (MRA-GNN), which employs Multi-granularity graph to achieve better generalization ability for few-shot learning. We first construct the Local Proposal Network (LPN) based on attention to generate local images from foreground images. The intra-cluster similarity and the inter-cluster dissimilarity are considered in the local images to generate discriminative features. Finally, we take the local images and original images as the input of multi-grained GNN models to perform classification. We evaluate our work by extensive comparisons with previous GNN approaches and other baseline methods on two benchmark datasets (i.e., miniImageNet and CUB). The experimental study on both of the supervised and semi-supervised few-shot image classification tasks demonstrates the proposed MRA-GNN significantly improves the performances and achieves the state-of-the-art results we know.
Read moreFew-shot learning for COVID-19 chest X-ray classification with imbalanced data: an inter vs. intra domain study
Medical image datasets are essential for training models used in computer-aided diagnosis, treatment planning, and medical research. However, some challenges are associated with these datasets, including variability in data distribution, data scarcity, and transfer learning issues when using models pre-trained from generic images. This work studies the effect of these challenges at the intra- and inter-domain level in few-shot learning scenarios with severe data imbalance. For this, we propose a methodology based on Siamese neural networks in which a series of techniques are integrated to mitigate the effects of data scarcity and distribution imbalance. Specifically, different initialization and data augmentation methods are analyzed, and four adaptations to Siamese networks of solutions to deal with imbalanced data are introduced, including data balancing and weighted loss, both separately and combined, and with a different balance of pairing ratios. Moreover, we also assess the inference process considering four classifiers, namely Histogram, kNN, SVM, and Random Forest. Evaluation is performed on three chest X-ray datasets with annotated cases of both positive and negative COVID-19 diagnoses. The accuracy of each technique proposed for the Siamese architecture is analyzed separately. The results are compared to those obtained using equivalent methods on a state-of-the-art CNN, achieving an average F1 improvement of up to 3.6%, and up to 5.6% of F1 for intra-domain cases. We conclude that the introduced techniques offer promising improvements over the baseline in almost all cases and that the technique selection may vary depending on the amount of data available and the level of imbalance.
Read more