• Home
  • Search
  • Spatio-temporal segmentation and object tracking
  • https://doi.org/10.5075/epfl-thesis-1618Copy DOI Icon

Spatio-temporal segmentation and object tracking

  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Visual information is taking a predominant place in our society. With the advent of digital technologies, visual information will find new applications in domains ranging from communication, commerce to entertainment. New functionalities will be required that permit an extended interaction with the visual information. In particular, this is the case for digital television and the related problem of video sequence compression. The extent of possible interactions depends on the manner in which the visual information is represented. Up to now, the canonical representation is used. Also referred to as the waveform representation, it is a technical artifact of the image capture procedure. This representation severely restricts the functionalities available to the end-user. The latter is unable to freely manipulate or customize the received visual information. In this dissertation, we propose to describe the visual information through a semantically meaningful representation. This representation directly derives from the scene content, which is decomposed in terms of its constituting objects. Consequently, the viewing process is totally disconnected from the image capture procedure. This permits full interactivity with the visual information, leading to enhanced functionalities for the end-user. The essence of this dissertation is to automatically define the objects forming the scene arid to automatically track them through the video sequence. Two main issues are identified: the initial segmentation of the objects and their tracking. The first issue deals with segmenting the scene into its constituent objects. This has to be performed on the basis of the information available in two consecutive frames. In order to solve the problem, a split-and-merge approach is used. First, the image is segmented into small, spatio-temporally homogeneous regions. This is achieved through a top-down approach where spatial, temporal and change information is combined. These regions are used as a starting point to define the objects forming the scene. A bottom-up approach is used which iteratively merges the regions. The propensity of the regions to form an object is assessed in terms of both spatial and temporal information. The second issue deals with segmenting and tracking the objects in the successive frames. The coherence between the successive segmentations is ensured by using past and current information. Also, the different objects composing the scene are identified throughout the video sequence. This identification relies on temporal, spatial and spatio-temporal features of the objects. The proposed representation of the visual information finds a natural application in video coding. A second generation video coding scheme is presented which combines compression efficiency with extended functionalities. In particular, scalable coding is achieved in terms of both scene content and object coding quality.

Similar Papers
  • Research Article
  • Citations9

Shadow segmentation and tracking in real-world conditions

  • Jan 01, 2004
  • Infoscience (Ecole Polytechnique Fédérale de Lausanne)
  • E Salvador
  • Supplementary Content

VLSI architectures design for encoders of High Efficiency Video Coding (HEVC) standard

  • Jan 01, 2016
  • Politecnico di Torino
  • Guoping Xiao
  • Research Article

Research on the Influence of Visual Illusion between Visually Perceived System and Visually Guided Action System

  • Jul 01, 2004
  • Chien-Hsiung Chen +1
  • Single Book
  • Citations200

Multidimensional Signal, Image, and Video Processing and Coding

  • Jan 01, 2012
  • John W Woods
  • Research Article

Selected topics on distributed video coding

  • Jan 01, 2008
  • Infoscience (Ecole Polytechnique Fédérale de Lausanne)
  • Mourad Ouaret
  • Book Chapter

Recent Advances in Computational Complexity Techniques for Video Coding Applications

  • Jan 01, 2013
  • Dan Grois +1
  • Conference Article
  • Citations1

Multi-scale Mobile Phone Playing Behavior Recognition Based on Temporal Enhancement and Interaction

  • Jun 18, 2021
  • Ming Fang +2
  • Conference Article
  • Citations9

3D Convolutional Neural network for Home Monitoring using Low Resolution Thermal-sensor Array

  • Jan 01, 2019
  • Lili Tao +4
  • Research Article

Rate-computation optimized block based video coding

  • Jan 01, 1998
  • Open Collections
  • Rabab K Ward +2
  • Research Article
  • Citations3

Effects of limitations on the use of some visual and kinaesthetic information in spatial mapping during exploration in the rat.

  • May 01, 1996
  • The Quarterly journal of experimental psychology. B, Comparative and physiological psychology
  • M C Buhot +3
  • Research Article
  • Citations155

Visual and phonological pathways to the lexicon: Evidence from Chinese readers

  • Jul 01, 1995
  • Memory & Cognition
  • K J Leck +2
  • Research Article

在IEEE 802.11e下針對視訊串流之動態優先分配

  • Jan 01, 2008
  • 張輔旺
  • Research Article
  • Citations5

An Efficient Predictive Watershed Video Segmentation Algorithm Using Motion Vectors

  • Mar 01, 2010
  • Journal of Information Science and Engineering
  • Kuo–Liang Chung +2
  • PDF
  • Research Article
  • Citations66

Visual and Non-Visual Contributions to the Perception of Object Motion during Self-Motion

  • Feb 07, 2013
  • PLoS ONE
  • Brett R Fajen +1
  • Supplementary Content

Do the eyes tell the truth? Mechanisms of peripheral-vision usage and practical implications

  • Jul 19, 2020
  • Bern Open Repository and Information System (University of Bern)
  • Christian Vater
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.