- Supplementary Content
- 10.3384/9789181182323
Spatiotemporal Learning for Motion Estimation and Visual Recognition
- Sep 15, 2025
- Yushan Zhang
The field of computer vision has undergone rapid development. Starting from recognition tasks such as classification, detection, and segmentation, the focus of visual analysis has gradually shifted towards learning spatiotemporal information. This thesis presents research on spatiotemporal learning, with a particular emphasis on motion estimation and visual recognition. First, we address the problem of video object tracking. Previous methods have primarily re-lied on learning improved appearance representations, while the spatiotemporal relationships of individual objects have been underexplored. We propose leveraging optical flow features to achieve higher generalization in semi-supervised video object segmentation, directly incorporating these features into both the target representation and the decoder network. Our experiments and analysis show that enriching feature representations with spatiotemporal information improves segmentation quality and generalization capability. Next, we investigate spatiotemporal learning in 3D for motion estimation, specifically scene flow estimation. Scene flow estimation as an important research topic in 3D computer vision is crucial for applications such as robotics, autonomous driving, embodied navigation, and tracking. We investigate the problem in different perspectives: 1. What is the best formulation for solving the problem and how to learn a better spatiotemporal feature representation? 2. Can we introduce uncertainty estimation to the task, which is of crucial importance for safety-critical downstream tasks? 3. How to scale the estimation to large-scale data, e.g., autonomous scenes, and leverage the temporal information without introducing much computation overheads? To answer these questions, we explore the use of transformers for improved feature representation, diffusion models for uncertainty estimation, and efficient feature learning methods for multi-frame, large-scale autonomous driving scenarios. Finally, we extend our research to joint visual segmentation, tracking, and open-vocabulary recognition in LiDAR sequences, particularly for autonomous scenes. In such environments, precise segmentation, tracking, and recognition of objects are essential for downstream analysis and control. Current human-annotated open-source datasets allow for reasonable tracking of traffic participants such as cars and pedestrians. However, we aim to advance beyond this towards segmenting and tracking any object in LiDAR data. To this end, we propose a pseudo-labeling engine that leverages the 2D visual foundation model SAM v2 and the vision-language model CLIP to automatically label LiDAR streams. We further introduce the SAL-4D model, capable of segmenting, tracking, and recognizing any object in a zero-shot manner. In summary, we explore the learning of spatiotemporal information in both 2D image and 3D point cloud domains. In the image domain, we demonstrate that spatiotemporal information improves video object segmentation quality and generalization. In the 3D point cloud domain, we show that spatiotemporal learning enables more accurate motion estimation and facilitates the first method for zero-shot segmentation, tracking, and open-vocabulary recognition of arbitrary objects.
Read more