- Research Article
1
- 10.1109/jiot.2025.3600817
M2FM: A Multimodal Fusion Model for Human Action Recognition With Camera and Millimeter-Wave Radar
- Nov 01, 2025
- IEEE Internet of Things Journal
- Jiangang Yi + 4 more +4
Human action recognition is a research hotspot in the field of ambient intelligence, serving as a foundation for the Internet of Healthcare Things (IoHT) and smart home wellness, with extensive application value. Currently, most research primarily focuses on the performance of single-modal approach for human action recognition. However, single-modal data cannot adequately capture the action characteristics of the human body. For instance, camera data lacks detailed micro-motion information, while radar data lacks visual appearance information. This limitation makes human action recognition systems susceptible to complex covariate influences. To address this issue, this paper proposes a Multi-Modal Fusion Model, M2FM, to accurately recognize various complex human action from millimeter-wave radar signals and video data. In the millimeter-wave radar branch, a LFNet network is constructed to capture richly hierarchical human action representation by extracting micro-Doppler feature and cadence velocity feature simultaneously. Specially, a Linear Feedforward Neural Network (LFNN) module is designed for modeling global-local feature of human action. In the video branch, a lightweight Transformer-based video analysis and action recognition network, STL-Former, is developed. Specially, an Agent downSampling (AS) attention module and a Linear Maxpooling (LM) attention module are designed to achieve context-aware downsampling with global receptive field and spatial vector sequence downsampling, respectively. Additionally, a multi-scale Linear attention (Litner) module is designed for modeling the spatio-temporal information of human action. Finally, the outputs from two branches are fused by a filter-based Dempster-Shafer theory. The experimental results on a self-collected Multi-modal Human Action Dataset, JH-MHAD, show that the M2FM model outperforms other advanced models, with an average recognition accuracy of 99.3% for eight different human actions, and it demonstrates strong adaptability under different conditions.
Read more