- Research Article
- 10.1080/01431161.2025.2579807
Multimodal fusion network with learnable wavelet-enhanced features for hyperspectral unmixing
- Oct 30, 2025
- International Journal of Remote Sensing
- Zhixiang Wang + 4 more +4
Hyperspectral unmixing (HU) aims to obtain subpixel-level material composition information, which is crucial for the precise advancement of hyperspectral image processing techniques. In recent years, deep learning (DL) has been widely applied to HU due to its strong ability to capture complex and nonlinear feature relationships in the data. However, relying solely on hyperspectral images for unmixing often fails to effectively distinguish objects in complex scenes, particularly when different endmembers exhibit similar spectral characteristics. To address this limitation, we propose a Multimodal Fusion Network (MMFNet) that incorporates the elevation information inherent in light detection and ranging (LiDAR) data. MMFNet is capable of simultaneously extracting spectral features from hyperspectral images and spatial features from LiDAR data. Furthermore, most existing DL-based HU methods operate only in the original spectral domain, which makes them susceptible to spectral variability, noise, and limited discriminative capacity. To overcome these challenges, we integrate Learnable Wavelet Transform (LWT) into MMFNet to adaptively decompose hyperspectral signals into multiple frequency subdomains, thereby mitigating spectral variability and noise while preserving spatial consistency. In addition, a Multi-Scale Convolution Fusion Module (MSCFM) is designed to capture semantic information at different receptive fields and enhance the fine-grained fusion of spectral and spatial features. Through these designs, MMFNet produces more robust and discriminative feature representations, enabling better separation of spectrally similar endmembers. Extensive experiments demonstrate that the proposed MMFNet outperforms several state-of-the-art unmixing methods.
Read more