- Research Article
3
- 10.1016/j.neucom.2019.04.024
Real-time object tracking via self-adaptive appearance modeling
- Apr 15, 2019
- Neurocomputing
- Ming Xin + 4 more +4
Real-time object tracking via self-adaptive appearance modeling
We present a novel local region approach for statistically characterizing appearance in the context of medical image segmentation via deformable models. Our appearance model reflects the inhomogeneity of tissue mixtures around the exterior of the object of interest by determining mixture-consistent local region types relative to the object boundary. The region types are formed by clustering local regional image descriptors. We partition the object boundary according to region type and apply principal component analysis on the cluster populations to acquire a statistical model of object appearance that accounts for local variability in the object exterior. We present results using this approach to segment bladders and prostates in CT in the context of day-to-day adaptive radiotherapy for prostate cancer. Results show improved fits versus those obtained with a previously developed method
Real-time object tracking via self-adaptive appearance modeling
Real-time object tracking via self-adaptive appearance modeling
Robust visual tracking based on generative and discriminative model collaboration
Effective object appearance model is one of the key issues for the success of visual tracking. Since the appearance of a target and the environment changes dynamically, the majority of existed visual tracking algorithms tend to drift away from targets. To address this issue, we propose a robust tracking algorithm by integrating the generative and discriminative model. The object appearance model is made up of generative target model and a discriminative classifier. For the generative target model, we adopt the weighted structural local sparse appearance model combining patch based gray value and Histogram of Oriented Gradients feature as the patch dictionary. By sampling positives and negatives, alignment-pooling features are obtained based on the patch dictionary through local sparse coding, then we use support vector machine to train the discriminative classifier. The proposed method is embedded into a Bayesian inference framework for visual tracking. A combined matching method is adopted to improve the proposal distribution of the particle filter. Moreover, in order to adapt the situation change, the patch dictionary and discriminative classifier are updated by incremental learning every five frames. Experimental results on some publicly available benchmarks of video sequences demonstrate the accuracy and effectiveness of our tracker.
Read moreOsteoporosis Presence Verification Using MACE Filter Based Statistical Models of Appearance with Application to Cervical X-ray Images
vertebral fracture is a very common outcome of osteoporosis, which is one of the major public health concerns in the world. Early detection of vertebral fractures is important because timely pharmacologic intervention can reduce the risk of subsequent additional fractures. Our goal seeks to develop a computerized method for detection of vertebral fractures by measuring the shape and appearance of vertebrae on cervical xray radiographs in order to assist radiologist’s image interpretation and thus allow the early diagnosis of osteoporosis. The statistical models of shape and appearance are powerful tools for interpreting medical images. This work introduces the application of correlation filter classifiers for identification and verification of the osteoporosis presence in cervical vertebrae training/ testing set. Correlation filter classifiers have been previously applied to other biometric classification tasks, but not to classification of cervical vertebrae images. We describe how the extraction of an appropriate region of interest in the cervical vertebrae surface can be used to design correlation filters that accomplish 90 % recognition on a database of 50 cervical bone shapes.Keywordsx-ray radiographsSegmentationASM modelingCorrelation filtervertebral deformityOsteoporosis
Read moreROAM: A Rich Object Appearance Model with Application to Rotoscoping
Rotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer the artists a much better interactive control on the definition, editing and manipulation of the segments of interest. Sticking to this prevalent rotoscoping paradigm, we propose a novel framework to capture and track the visual aspect of an arbitrary object in a scene, given a first closed outline of this object. This model combines a collection of local foreground/background appearance models spread along the outline, a global appearance model of the enclosed object and a set of distinctive foreground landmarks. The structure of this rich appearance model allows simple initialization, efficient iterative optimization with exact minimization at each step, and on-line adaptation in videos. We demonstrate qualitatively and quantitatively the merit of this framework through comparisons with tools based on either dynamic segmentation with a closed curve or pixel-wise binary labelling.
Read moreROAM: A Rich Object Appearance Model with Application to Rotoscoping.
Rotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer the artists a much better interactive control on the definition, editing and manipulation of the segments of interest. Sticking to this prevalent rotoscoping paradigm, we propose a novel framework to capture and track the visual aspect of an arbitrary object in a scene, given an initial closed outline of this object. This model combines a collection of local foreground/background appearance models spread along the outline, a global appearance model of the enclosed object and a set of distinctive foreground landmarks. The structure of this rich appearance model allows simple initialization, efficient iterative optimization with exact minimization at each step, and on-line adaptation in videos. We further extend this model by so-called trimaps which serve as an input to alpha-matting algorithms to allow truly seamless compositing. To this end, we leverage local classifiers attached to the roto-curves to define a confidence measure that is well-suited to define trimaps with adaptive band-widths. The resulting trimaps are parametric, temporally consistent and remain fully editable by the artist. We demonstrate qualitatively and quantitatively the merit of this framework through comparisons with tools based on either dynamic segmentation with a closed curve or pixel-wise binary labelling.
Read moreShape Priors and Online Appearance Learning for Variational Segmentation and Object Recognition in Static Scenes
We present an integrated two-level approach to computationally analyzing image sequences of static scenes by variational segmentation. At the top level, estimated models of object appearance and background are probabilistically fused to obtain an a-posteriori probability for the occupancy of each pixel. The data-association strategy handles object occlusions explicitly. At the lower level, object models are inferred by variational segmentation based on image data and statistical shape priors. The use of shape priors allows to distinguish between recognition of known objects and segmentation of unknown objects. The object models are sufficiently flexible to enable the integration of general cues like advanced shape distances. At the same time, they are highly constrained from the optimization viewpoint: the globally optimal parameters can be computed at each time instant by dynamic programming. The novelty of our approach is the integration of state-of-the-art variational segmentation into a probabilistic framework for static scene analysis that combines both on-line learning and prior knowledge of various aspects of object appearance.
Read moreLearning Local–Global Multiple Correlation Filters for Robust Visual Tracking with Kalman Filter Redetection
Visual object tracking is a significant technology for camera-based sensor networks applications. Multilayer convolutional features comprehensively used in correlation filter (CF)-based tracking algorithms have achieved excellent performance. However, there are tracking failures in some challenging situations because ordinary features are not able to well represent the object appearance variations and the correlation filters are updated irrationally. In this paper, we propose a local–global multiple correlation filters (LGCF) tracking algorithm for edge computing systems capturing moving targets, such as vehicles and pedestrians. First, we construct a global correlation filter model with deep convolutional features, and choose horizontal or vertical division according to the aspect ratio to build two local filters with hand-crafted features. Then, we propose a local–global collaborative strategy to exchange information between local and global correlation filters. This strategy can avoid the wrong learning of the object appearance model. Finally, we propose a time-space peak to sidelobe ratio (TSPSR) to evaluate the stability of the current CF. When the estimated results of the current CF are not reliable, the Kalman filter redetection (KFR) model would be enabled to recapture the object. The experimental results show that our presented algorithm achieves better performances on OTB-2013 and OTB-2015 compared with the other latest 12 tracking algorithms. Moreover, our algorithm handles various challenges in object tracking well.
Read moreModel-free tracker for multiple objects using joint appearance and motion inference.
Model-free tracking is a widely-accepted approach to track an arbitrary object in a video using a single frame annotation with no further prior knowledge about the object of interest. Extending this problem to track multiple objects is really challenging because: a) the tracker is not aware of the objects' type while trying to distinguish them from background (detection task), and b) The tracker needs to distinguish one object from other potentially similar objects (data association task) to generate stable trajectories. In order to track multiple arbitrary objects, most existing model-free tracking approaches rely on tracking each target individually by updating their appearance model independently. Therefore, in this scenario they often fail to perform well due to confusion between the appearance of similar objects, their sudden appearance changes and occlusion. To tackle this problem, we propose to use both appearance and motion models, and to learn them jointly using graphical models and deep neural networks features. We introduce an indicator variable to predict sudden appearance change and/or occlusion. When these happen, our model does not update the appearance model thus avoiding using the background and/or incorrect object to update the appearance of the object of interest mistakenly, and relies on our motion model to track. Moreover, we consider the correlation among all targets, and seek the joint optimal locations for all targets simultaneously as a graphical model inference problem. We learn the joint parameters for both appearance model and motion model in an online fashion under the framework of LaRank. Experiment results show that our method achieved superior performance compared to the competitive methods.
Read moreAdaptive particle filters for visual object tracking using joint PCA appearance model and consensus point correspondences
This paper addresses issues on moving object tracking from videos. We propose a novel tracking scheme that jointly exploits local object features using consensus point correspondences, and global object appearance and shape models using adaptive particle filter-based eigen-tracking. The paper include the following main novelties: (a) employ consensus feature point correspondences to estimate the motion vector of shape model; (b) employ adaptive particle filters and motioncorrected state vector for joint appearance- and shape-based eigen-tracking. An adaptive number of particles is chosen automatically based on an updated estimation of covariancematrix. Further, online learning is made adaptive to avoid learning using partially-occluded objects. The proposed scheme is realized by integrating SURF and RANSAC [8, 9] for estimating consensus point correspondences, and modify an existing particle filter-based eigen-tracking [4]. Experiment results on tracking moving objects in videos have shown that the proposed scheme provides more accurate tracking, especially for objects with fast motion or long-term partial occlusions. The average number of particles is significantly reduced. Comparisons have been made with an existing method, results have shown that the proposed scheme has provided an improved tracking accuracy at the cost of more computations.
Read moreModeling Self-Occlusions/Disocclusions in Dynamic Shape and Appearance Tracking for Obtaining Precise Shape
Modeling Self-Occlusions/Disocclusions in Dynamic Shape and Appearance Tracking for Obtaining Precise Shape We present a method to determine the precise shape of a dynamic object from video. This problem is fundamental to computer vision, and has a number of applications, for example, 3D video/cinema post-production, activity recognition and augmented reality. Current tracking algorithms that determine precise shape can be roughly divided into two categories: 1) Global statistics partitioning methods, where the shape of the object is determined by discriminating global image statistics, and 2) Joint shape and appearance matching methods, where a template of the object from the previous frame is matched to the next image. The former is limited in cases of complex object appearance and cluttered background, where global statistics cannot distinguish between the object and background. The latter is able to cope with complex appearance and a cluttered background, but is limited in cases of camera viewpoint change and object articulation, which induce self-occlusions and self-disocclusions of the object of interest. The purpose of this thesis is to model self-occlusion/disocclusion phenomena in a joint shape and appearance tracking framework. We derive a non-linear dynamic model of the object shape and appearance taking into account occlusion phenomena, which is then used to infer self-occlusions/disocclusions, shape and appearance of the object in a variational optimization framework. To ensure robustness to other unmodeled phenomena 5 that are present in real-video sequences, the Kalman filter is used for appearance updating. Experiments show that our method, which incorporates the modeling of self-occlusion/disocclusion, increases the accuracy of shape estimation in situations of viewpoint change and articulation, and out-performs current state-of-the-art methods for shape tracking.
Read morePyramidCore – Feature Pyramids for Few-Shot Logical Anomaly Detection
Recent few-shot logical anomaly detection methods rely on external information for accurate detection. This is often done through handmade text prompts and category-specific procedures, making them infeasible to apply to new datasets. Full-shot methods do not utilise this additional information but extract meaningful representations of local and global structures. We hypothesise that a major drawback of few-shot logical anomaly detection methods is the over-reliance on external information and suboptimal image representation. However, matching the representations learned by full-shot methods is challenging due to the lack of data in a few-shot setting. We propose PyramidCore, a novel few-shot logical anomaly detection method that does not rely on external information but instead uses a robust appearance model that can be built from only a few examples. It builds a hierarchical model of object appearance, enabling the detection of complex logical anomalies at different scales. The proposed method achieves state-of-the-art results on the challenging MVTec LOCO Dataset.
Read moreNotice of Removal: The Visual Object Tracking VOT2013 Challenge Results
Visual tracking has attracted a significant attention in the last few decades. The recent surge in the number of publications on tracking-related problems have made it almost impossible to follow the developments in the field. One of the reasons is that there is a lack of commonly accepted annotated data-sets and standardized evaluation protocols that would allow objective comparison of different tracking methods. To address this issue, the Visual Object Tracking (VOT) workshop was organized in conjunction with ICCV2013. Researchers from academia as well as industry were invited to participate in the first VOT2013 challenge which aimed at single-object visual trackers that do not apply pre-learned models of object appearance (model-free). Presented here is the VOT2013 benchmark dataset for evaluation of single-object visual trackers as well as the results obtained by the trackers competing in the challenge. In contrast to related attempts in tracker benchmarking, the dataset is labeled per-frame by visual attributes that indicate occlusion, illumination change, motion change, size change and camera motion, offering a more systematic comparison of the trackers. Furthermore, we have designed an automated system for performing and evaluating the experiments. We present the evaluation protocol of the VOT2013 challenge and the results of a comparison of 27 trackers on the benchmark dataset. The dataset, the evaluation tools and the tracker rankings are publicly available from the challenge website (http://votchallenge.net).
Read moreRemote Sensing Object Tracking With Deep Reinforcement Learning Under Occlusion
Object tracking is an important research direction of space Earth observation in the field of remote sensing. Although the existing correlation filter-based and deep learning (DL)-based object tracking algorithms have achieved great success, they are still unsatisfactory for the problem of object occlusion. The occlusion caused by the complex change in background, and the deviation of the tracking lens, causes object information to go missing, which leads to the omission of detection. Traditionally, most methods for object tracking under occlusion adopt a complex network model, which redetects the occluded object. To address this issue, we propose a novel object tracking approach. First, an action decision-occlusion handling network (AD-OHNet) based on deep reinforcement learning (DRL) is built to achieve low computational complexity for object tracking under occlusion. Second, the temporal and spatial context, the object appearance model, and the motion vector are adopted to provide the occlusion information, which drives actions in reinforcement learning under complete occlusion and contributes to improving the accuracy of tracking while maintaining speed. Finally, the proposed AD-OHNet is evaluated on three remote sensing video datasets of Bogota, Hong Kong, and San Diego taken from Jilin-1 commercial remote sensing satellites. The video datasets all shared problems of low spatial resolution, background clutter, and small objects. Experimental results on the three video datasets validate the effectiveness and efficiency of the proposed tracker.
Read moreHead Tracking with Shape Modeling and Detection
Color-based tracking has proved efficient and robust recently. Trackers build the object appearance model with histogram statistics, search and evaluate hypothesis in a probabilistic framework. This method relies much on the discrimination between object and scene blobs. Color clutter in the scene, although not so many in quantity, may distract these trackers. We build explicitly object shape model and insert the head detector into the observation model to resist these clutters in the scene for improved tracker. The detector scans the image and output probability value as the possibility of current window being a candidate human head. Experiments demonstrate the method can work more accurately and robustly.
Read moreStructural sparse representation-based semi-supervised learning and edge detection proposal for visual tracking
In discriminative tracking, lots of tracking methods easily suffer from changes of pose, illumination and occlusion. To deal with this problem, we propose a novel object tracking method using structural sparse representation-based semi-supervised learning and edge detection. First, the object appearance model is constructed by extracting sparse code features on different layers to exploit local information and holistic information. To utilize unlabelled samples information, the semi-supervised learning is introduced and a classifier is trained which is used to measure candidates. In addition, an auxiliary positive sample set is maintained to improve the performance of the classifier. We subsequently adopt an edge detection to alleviate the error accumulation based on the ranking results from the learned classifier. Finally, the proposed method is implemented under the Bayesian inference framework. Both the proposed tracker and several current trackers are tested on some challenging videos, where the target objects undergo pose change, illumination and occlusion. The experimental results demonstrate that the proposed tracker outperforms the other state-of-the-art methods in terms of effectiveness and robustness.
Read more