- Research Article
- 10.70286/isu-25.03.2026.013
FORMATS OF HUMAN MOTION DATA AND THEIR SUITABILITY FOR SPATIO-TEMPORAL DEEP LEARNING ANALYSIS
- Mar 25, 2026
- Artur Bielobrov + 1 more +1
The rapid proliferation of human-computer interaction (HCI) systems, virtual reality (VR), augmented reality (AR), and contactless interfaces has intensified the demand for reliable automated recognition and interpretation of human gestures.A fundamental prerequisite for constructing accurate spatio-temporal gesture recognition models lies in the choice of input data format, as each representation captures different aspects of spatial structure, temporal dynamics, and motion semantics.In this context, the present paper focuses on four principal formats of human motion data used in gesture recognition-RGB video, optical flow, skeleton (pose) models, and keypoint representations-and analyzes how their properties influence downstream deep learning pipelines.Specifically, the study investigates how these formats interact with state-of-the-art spatio-temporal architectures and how they affect performance on a standard benchmark.To this end, we evaluate the suitability of each data format for spatio-temporal analysis within deep learning architectures by comparing their behavior when coupled with representative models on the Jester dataset.Modern deep learning approaches, including 3D Convolutional Neural Networks (3D CNN [1]), hybrid CNN+LSTM architectures, and Transformer-based models [2][3][4], each impose distinct requirements on the input data representation, which in turn affects classification performance and inference speed.Therefore, this work systematically relates specific motion data formats to concrete model families, examining how choices such as RGB versus optical flow or skeleton versus keypoints shape achievable accuracy, latency, and robustness on Jester.A structured comparison of data formats, grounded in empirical results on a common benchmark, is thus essential for selecting an optimal configuration of input representation and model architecture for real-world deployment scenarios.RGB video is the most accessible and widely used data format for gesture recognition.It captures full colour information of each frame in a video sequence and preserves the complete visual context of the scene.RGB video serves as the primary input for architectures such as 3D CNN, which process the data as a spatio-temporal tensor of dimensions (height width time), simultaneously extracting spatial features and temporal patterns through three-dimensional convolutional kernels.When applied
Read more