- Conference Article
- 10.1109/iccvw69036.2025.00018
Aligning Multimodal Data for Fine-Grained Video Understanding via Cross-Attentive Recurrent Fusion
- Oct 19, 2025
- Nam-Ho Kim + 1 more +1
Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text representations using GRU-based sequence encoders and cross-modal attention mechanisms. The model is trained using a combination of classification or regression loss, depending on the task, and is further regularized through feature-level augmentation and autoencoding techniques. To evaluate the generality of our framework, we conduct experiments on two challenging benchmarks: the DVD dataset for real-world Violence Detection and the Aff-Wild2 dataset for Valence-Arousal estimation. Our results demonstrate that the proposed fusion strategy significantly outperforms unimodal baselines, with cross-attention and feature augmentation contributing notably to robustness and performance. The code used in the 9th ABAW competition is available at https://github.com/namho-96/ABAW-9th.
Read more