- Research Article
6
- 10.1109/tgrs.2023.3330568
Hypothesis Margin-Based Ensemble Method for the Classification of Noisy Remote Sensing Data
- Jan 01, 2023
- IEEE Transactions on Geoscience and Remote Sensing
- Wei Feng + 6 more +6
The accuracy of a classifier, whether it is an ensemble or not, is directly influenced by the training data used in learning. In remote sensing, training data mislabeling is inevitable and faces a major challenge. This paper proposes a versatile data cleaning which handles the mislabeling problem by exploiting the ensemble concepts for identifying, then eliminating or correcting the mislabeled training data. A powerful ensemble method, random forest, is at the core of our filter design and helps to distinguish mislabeled data from uncorrupted data more accurately. The major contribution of this work lies on the explicit use of the hypothesis margin as a decision means to identify and eliminate or correct mislabeled training data in an ensemble learning framework. Another key development that makes our algorithm superior to existing approaches is a design that avoids rare class instances to be mistaken for class noise. This fundamental aspect makes our data cleaning system particularly suitable for remote sensing classification tasks which usually suffer from both mislabeling and imbalance problems. The effectiveness of our algorithm is demonstrated in performing mapping of land covers. The generalization performance of two major supervised noise-sensitive classifiers, boosting and K-nearest neighbors, is strengthened by effective class noise reduction. A comparative analysis is conducted with respect to random forest, deep convolutional neural networks, as well as two well-established ensemble-based class noise filters, the majority vote and the consensus vote filters. This analysis demonstrates that our approach is more accurate than deep convolutional neural networks (one-dimensional CNN, AlexNet, EfficientNet, ResNet50 and ShuffletNet) and the reference ensemble methods.
Read more