- Research Article
- 10.1016/j.jii.2025.101020
Emerging perspectives on embodied intelligence in future smart manufacturing
- Mar 01, 2026
- Journal of Industrial Information Integration
- Dan Xia + 3 more +3
Publications from 2021 to 2026
Showing 5 of 5 papers
Emerging perspectives on embodied intelligence in future smart manufacturing
Semantic SLAM for dynamic semi-structured environments using object and plane detection
Context Sensing Attention Network for Video-based Person Re-identification
Video-based person re-identification (ReID) is challenging due to the presence of various interferences in video frames. Recent approaches handle this problem using temporal aggregation strategies. In this work, we propose a novel Context Sensing Attention Network (CSA-Net), which improves both the frame feature extraction and temporal aggregation steps. First, we introduce the Context Sensing Channel Attention (CSCA) module, which emphasizes responses from informative channels for each frame. These informative channels are identified with reference not only to each individual frame, but also to the content of the entire sequence. Therefore, CSCA explores both the individuality of each frame and the global context of the sequence. Second, we propose the Contrastive Feature Aggregation (CFA) module, which predicts frame weights for temporal aggregation. Here, the weight for each frame is determined in a contrastive manner: i.e., not only by the quality of each individual frame, but also by the average quality of the other frames in a sequence. Therefore, it effectively promotes the contribution of relatively good frames. Extensive experimental results on four datasets show that CSA-Net consistently achieves state-of-the-art performance.
Read moreSpeechFormer: A Hierarchical Efficient Framework Incorporating the Characteristics of Speech
Transformer has obtained promising results on cognitive speech signal processing field, which is of interest in various applications ranging from emotion to neurocognitive disorder analysis. However, most works treat speech signal as a whole, leading to the neglect of the pronunciation structure that is unique to speech and reflects the cognitive process. Meanwhile, Transformer has heavy computational burden due to its full attention operation. In this paper, a hierarchical efficient framework, called SpeechFormer, which considers the structural characteristics of speech, is proposed and can be served as a general-purpose backbone for cognitive speech signal processing. The proposed SpeechFormer consists of frame, phoneme, word and utterance stages in succession, each performing a neighboring attention according to the structural pattern of speech with high computational efficiency. SpeechFormer is evaluated on speech emotion recognition (IEMOCAP & MELD) and neurocognitive disorder detection (Pitt & DAIC-WOZ) tasks, and the results show that SpeechFormer outperforms the standard Transformer-based framework while greatly reducing the computational cost. Furthermore, our SpeechFormer achieves comparable results to the state-of-the-art approaches.
Read more<title>Robust exterior autonomous navigation</title>
This paper introduces a proposed approach for robust exterior autonomous navigation, based on the use of redundant, and complimentary position estimation sensors, fused and integrated to provide robust operation. The proposed approach considers constraints due to a combination of theoretical and practical limitations of the sensors performance characteristics, and sensor fusion and integration algorithm limitations. A hypothetical exterior autonomous navigation application is also introduced to illustrate a systems perspective on one class of autonomous navigation, specifically routine autonomous navigation in a known, semi-structured environment. The focus of this paper is primarily on the autonomous vehicles position estimation sensor suite selection, and sensor fusion and integration approach.
Read more