• Home
  • Search
  • Spatiotemporal Learning for Motion Estimation and Visual Recognition
  • https://doi.org/10.3384/9789181182323Copy DOI Icon

Spatiotemporal Learning for Motion Estimation and Visual Recognition

  • Sep 15, 2025
  • Yushan Zhang
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

The field of computer vision has undergone rapid development. Starting from recognition tasks such as classification, detection, and segmentation, the focus of visual analysis has gradually shifted towards learning spatiotemporal information. This thesis presents research on spatiotemporal learning, with a particular emphasis on motion estimation and visual recognition. First, we address the problem of video object tracking. Previous methods have primarily re-lied on learning improved appearance representations, while the spatiotemporal relationships of individual objects have been underexplored. We propose leveraging optical flow features to achieve higher generalization in semi-supervised video object segmentation, directly incorporating these features into both the target representation and the decoder network. Our experiments and analysis show that enriching feature representations with spatiotemporal information improves segmentation quality and generalization capability. Next, we investigate spatiotemporal learning in 3D for motion estimation, specifically scene flow estimation. Scene flow estimation as an important research topic in 3D computer vision is crucial for applications such as robotics, autonomous driving, embodied navigation, and tracking. We investigate the problem in different perspectives: 1. What is the best formulation for solving the problem and how to learn a better spatiotemporal feature representation? 2. Can we introduce uncertainty estimation to the task, which is of crucial importance for safety-critical downstream tasks? 3. How to scale the estimation to large-scale data, e.g., autonomous scenes, and leverage the temporal information without introducing much computation overheads? To answer these questions, we explore the use of transformers for improved feature representation, diffusion models for uncertainty estimation, and efficient feature learning methods for multi-frame, large-scale autonomous driving scenarios. Finally, we extend our research to joint visual segmentation, tracking, and open-vocabulary recognition in LiDAR sequences, particularly for autonomous scenes. In such environments, precise segmentation, tracking, and recognition of objects are essential for downstream analysis and control. Current human-annotated open-source datasets allow for reasonable tracking of traffic participants such as cars and pedestrians. However, we aim to advance beyond this towards segmenting and tracking any object in LiDAR data. To this end, we propose a pseudo-labeling engine that leverages the 2D visual foundation model SAM v2 and the vision-language model CLIP to automatically label LiDAR streams. We further introduce the SAL-4D model, capable of segmenting, tracking, and recognizing any object in a zero-shot manner. In summary, we explore the learning of spatiotemporal information in both 2D image and 3D point cloud domains. In the image domain, we demonstrate that spatiotemporal information improves video object segmentation quality and generalization. In the 3D point cloud domain, we show that spatiotemporal learning enables more accurate motion estimation and facilitates the first method for zero-shot segmentation, tracking, and open-vocabulary recognition of arbitrary objects.

Similar Papers
  • Research Article

MoBox: Enhancing Video Object Segmentation With Motion-Augmented Box Supervision

  • Jan 01, 2025
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Xiaomin Li +6
  • Dissertation

Investigation on Segmentation, Recognition and 3D Reconstruction of Objects Based on LiDAR Data Or MRI

  • May 01, 2015
  • Shijun Tang
  • Research Article

Semi-Supervised Video Object Segmentation with Global Feature Enhancement and Mask Correction

  • Jul 01, 2024
  • Journal of Computer-Aided Design & Computer Graphics
  • Zuwang Pan +3
  • Research Article
  • Citations5

Semi-Supervised Video Object Segmentation Based on Local and Global Consistency Learning

  • Jan 01, 2021
  • IEEE Access
  • Huagang Liang +3
  • Conference Article
  • Citations24

Pixel-Level Bijective Matching for Video Object Segmentation

  • Jan 01, 2022
  • Suhwan Cho +4
  • Conference Article
  • Citations22

Towards Robust Video Object Segmentation with Adaptive Object Calibration

  • Oct 10, 2022
  • Xiaohao Xu +3
  • Research Article
  • Citations23

Object recognition and segmentation in videos by connecting heterogeneous visual features

  • Feb 01, 2008
  • Computer Vision and Image Understanding
  • Valérie Gouet-Brunet +1
  • Conference Article
  • Citations1

Hierarchical Embedding Guided Network for Video Object Segmentation

  • Sep 19, 2021
  • Chin-Hsuan Shih +1
  • Conference Article
  • Citations2

Fast motion estimation algorithm using dual bit-plane matching criteria

  • May 01, 2014
  • Changryoul Choi +1
  • Research Article
  • Citations244

Neurocomputational bases of object and face recognition.

  • Aug 29, 1997
  • Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences
  • Irving Biederman +1
  • Conference Article
  • Citations5

On the complexity and accuracy of motion estimation -using lie operators

  • Sep 27, 2004
  • M Nalasani +1
  • Research Article
  • Citations3

Partial closed‐loop versus open‐loop motion estimation for HDTV compression

  • Dec 01, 1994
  • International Journal of Imaging Systems and Technology
  • Jeffrey S Mcveigh +1
  • Research Article
  • Citations2

Augmented Reality Frameworks for Object Recognition in Learning Application Domain: A Systematic Review

  • Sep 07, 2023
  • Journal of Advanced Research in Applied Sciences and Engineering Technology
  • Intan Nadiah Abdul Hakim +1
  • Book Chapter
  • Citations6

Probabilistic Search for Object Segmentation and Recognition

  • Jan 01, 2002
  • Ulrich Hillenbrand +1
  • Research Article
  • Citations39

Object and anatomical feature recognition in surgical video images based on a convolutional neural network

  • Jan 01, 2021
  • International Journal of Computer Assisted Radiology and Surgery
  • Yoshiko Bamba +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.