- Conference Article
- 10.1109/ieeeconf67917.2025.11443703
Multi-Modal Transformer Network for Classroom Instruction Analysis
- Oct 26, 2025
- Matthew Korban + 3 more +3
We propose a multimodal transformer for automatic classroom activity recognition that combines a background-suppressed Vision Transformer video model with a BERT-based transcript encoder. Each modality first predicts activity labels independently, and a cross-modal transformer then refines video predictions using aligned transcript cues. Evaluated on 50 hours of real-world K–12 classroom video and audio, our framework accurately classifies activities such as lecture, individual work, and group work, and consistently outperforms a strong video-only baseline. Per-frame F1 and cross-modal correlation analyses demonstrate that transcript signals substantially boost video recognition in authentic classrooms, highlighting the value of multimodal context for educational video analytics.
Read more