- Research Article
1017
- 10.1016/j.artint.2020.103448
Multiple object tracking: A literature review
- Dec 30, 2020
- Artificial Intelligence
- Wenhan Luo + 5 more +5
Multiple object tracking: A literature review
Multiple Object Tracking (MOT) has attracted increasing attention due to its academic and commercial potential. However, it still remains challenging due to factors like data association ambiguity and abrupt appearance changes. To address these issues, we extended the multiple hypothesis tracking (MHT) by incorporating the appearance features which extracted from faster region convolutional neural network (Faster-RCNN). Additionally, a principal component analysis method was introduced to update discriminative appearance features. Thus, the data association ambiguity can be reduced surprisingly by exploiting Faster-RCNN, which means fewer hypothesis for MHT. Meanwhile, the issue of abrupt appearance changes can be addressed by modeling discriminative appearance feature based-on MHT. Many experiments on MOT benchmark show competitive performance of the proposed approach in comparison with other state-of-art tracking methods.
Multiple object tracking: A literature review
Multiple object tracking: A literature review
Lightweight and Deep Appearance Embedding for Multiple Object Tracking
The main challenge of Multiple Object Tracking (MOT) is that there is great uncertainty in data association when using the tracked predicted values and tracked trajectories. Meanwhile, the MOT is complex and time‐consuming. When equipment resources are limited or real‐time requirement is high, its application is very limited. Therefore, we propose a Lightweight Deep Appearance Embedding (LDAE) to assist the association of trajectories. Firstly, in addition to motion information in data association, we also introduce more discriminative appearance features to participate in the affinity measure to effectively distinguish similar targets. Secondly, according to the idea of feature mapping, we design a lightweight deep appearance embedding module. It can help extract appearance features with less computation. Finally, we propose a simulated occlusion strategy for the training of the LDAE, which helps improve the ability to recognise different targets in dense scenes. The LDAE dramatically reduces the computational cost and improves the accuracy of data association. Extensive experiments are conducted on the MOT datasets (MOT16, MOT17 and MOT20), which prove that the LDAE outperforms several state‐of‐the‐art trackers in the tracking accuracy and anti‐occlusion performance. Furthermore, we apply the LDAE to escalators, which can achieve fast and stable tracking effect.
Read moreOnline Appearance-Motion Coupling for Multi-Person Tracking in Videos
Multi-person tracking in videos is a promising but challenging visual task. Recent progress in this field is introducing deep convolutional features as appearance models, which achieves robust tracking results when coupled with proper motion models. However, model failures that often cause severe tracking problems, have not been well discussed and addressed in previous work. In this paper, we propose a solution by online detecting such failures and accordingly adjusting the coupling between appearance and motion models. The strategy is letting the functional models take over when certain model faces data association ambiguity, and at the same time suppressing the influence of inappropriate observations during model update. Experimental results prove the benefit of our proposed improvement. Multiple object tracking; deep neural network; online learning; tracking-by-detection; multiple hypothesis tracking (key words)
Read moreMultiple Maneuvering Target Tracking Using MHT and Nonlinear Non-Gaussian Kalman Filter
In this paper, an algorithm for tracking multiple maneuvering targets by Multiple Hypothesis Tracking (MHT) with nonlinear non-Gaussian Kalman filter is investigated. The main challenges in multiple maneuvering targets tracking are the nonlinearity and non -Gaussianity problems. The Multiple Hypothesis Tracking (MHT) is used to detect the multiple targets in maneuverable and non-maneuverable modes. The computational requirements increase exponentially with number of tracks, the backscan depth and this can be reduced by careful design and tuning of MHT. The 1-backscan MHT algorithm is a good compromise between the two conflicting requirements of good tracking performance and limitation of computation time. The nonlinear non-Gaussian Kalman filter is used to track the target with high maneuver rate. The nonlinear non-Gaussian Kalman filter is implemented in MHT to give less probability of missing the target. The 1-backscan MHT with nonlinear non-Gaussian Kalman filter is free from computational burden by using simple probability concepts. This method of tracking also shows the reduction in the overshoot of root mean square error (RMSE).
Read moreRobust tracking and event detection using receding horizon incremental locally linear embeddings
This thesis addresses the problem of tracking under severe appearance changes using a combination of receding horizon and Incremental Locally Linear Embeddings (ILLE) techniques that provides adaptation to both time varying dynamics and abrupt appearance changes. In addition to enabling sustained tracking under these conditions, the proposed technique allows for efficiently detecting appearance changes by simply computing the distance between successive embeddings. These results are illustrated with several examples showing the ability of the proposed technique to flag appearance changes while not breaking track.
Read moreDisaster damage detection through synergistic use of deep learning and 3D point cloud features derived from very high resolution oblique aerial images, and multiple-kernel-learning
Disaster damage detection through synergistic use of deep learning and 3D point cloud features derived from very high resolution oblique aerial images, and multiple-kernel-learning
Read morePerson re-identification based on multi-appearance model
Person re-identification plays important roles in many practical applications. Due to various human poses, complex backgrounds and similarity of person clothes, person re-identification is still a challenging task. In this paper, we mainly focus on the robust and discriminative appearance feature representation and proposed a novel multi-appearance method for person re-identification. First, we proposed a deep feature fusion method and get the multi-appearance feature by combining two Convolutional Neural Networks. Then, in order to further enhance the representation of the appearance feature, the multi-part model was constructed by combining the whole body and the six body parts. Additionally, we optimized the feature extraction process by adding a pooling layer. Comprehensive and comparative experiments with the state-of-the-art methods over publicly available datasets demonstrated that the proposed method can get promising results.
Read more基于深度学习的目标跟踪方法研究现状与展望
The inverse synthetic aperture lidar (ISAL) have attracted increasing attention for its merits including small visual tracking which is considered as one of the important research topics in the field of computer vision due to its key role in versatile applications, such as precision guidance, intelligent video surveillance, human-computer interaction, robot navigation and public safety. The basic idea for implementing visual tracking is composed of finding the target object in a video or sequence of images, then determining its exact position in the next successive frames and finally generating the corresponding trajectory of this object. Visual tracking, however, is still a challenging problem in practice while taking into account the abrupt appearance changes of the target objects induced by their non-rigid transformation, the sophisticated lighting variation, the obstruction by the block or similar objects in the background and the camera jitter. Motivated by the successful applications in target detection and recognition in recent years, plenty of deep learning models have been integrated in the visual tracking and better performance over traditional methods was achieved in a series of data evaluations, which opens a new door in the field of visual tracking. In this paper, the overview and progress on visual tracking were summarized. The current challenges and corresponding solving approaches in this field are introduced firstly and in particular, several novel and mainstream visual tracking algorithms based on the deep learning are specially described and analyzed in details, including their basic ideas, advantages and disadvantages and future prospect.
Read moreMultiple hypothesis tracking based on the Shiryayev sequential probability ratio test
To date, Wald sequential probability ratio test (WSPRT) has been widely applied to track management of multiple hypothesis tracking (MHT). But in a real situation, if the false alarm spatial density is much larger than the new target spatial density, the original track score will be very close to the deletion threshold of the WSPRT. Consequently, all tracks, including target tracks, may easily be deleted, which means that the tracking performance is sensitive to the tracking environment. Meanwhile, if a target exists for a long time, its track will have a high score, which will make the track survive for a long time even after the target has disappeared. In this paper, to consider the relationship between the hypotheses of the test, we adopt the Shiryayev SPRT (SSPRT) for track management in MHT. By introducing a hypothesis transition probability, the original track score can increase faster, which solves the first problem. In addition, by setting an independent SSPRT for track deletion, the track score can decrease faster, which solves the second problem. The simulation results show that the proposed SSPRT-based MHT can achieve better tracking performance than MHT based on the WSPRT under a high false alarm spatial density.
Read moreMedical Image Classification Algorithm Based on Visual Attention Mechanism-MCNN.
Due to the complexity of medical images, traditional medical image classification methods have been unable to meet the actual application needs. In recent years, the rapid development of deep learning theory has provided a technical approach for solving medical image classification. However, deep learning has the following problems in the application of medical image classification. First, it is impossible to construct a deep learning model with excellent performance according to the characteristics of medical images. Second, the current deep learning network structure and training strategies are less adaptable to medical images. Therefore, this paper first introduces the visual attention mechanism into the deep learning model so that the information can be extracted more effectively according to the problem of medical images, and the reasoning is realized at a finer granularity. It can increase the interpretability of the model. Additionally, to solve the problem of matching the deep learning network structure and training strategy to medical images, this paper will construct a novel multiscale convolutional neural network model that can automatically extract high-level discriminative appearance features from the original image, and the loss function uses the Mahalanobis distance optimization model to obtain a better training strategy, which can improve the robust performance of the network model. The medical image classification task is completed by the above method. Based on the above ideas, this paper proposes a medical classification algorithm based on a visual attention mechanism-multiscale convolutional neural network. The lung nodules and breast cancer images were classified by the method in this paper. The experimental results show that the accuracy of medical image classification in this paper is not only higher than that of traditional machine learning methods but also improved compared with other deep learning methods, and the method has good stability and robustness.
Read moreThe Effects of Visual and Auditory Dual-task on Multiple Object Tracking Performance: Interference or Promotion?
Multiple Object Tracking(MOT) developed by Pylyshyn and Storm(1988) has been widely used in the study of capacity-limited and object-based attention. Many researchers interested in finding out whether or not the tracking processes in MOT occupy attentional resource and what type of resource is used. The typical paradigm used in this line of research is dual-task paradigm. Participants were asked to perform a visual or auditory task and the MOT task simultaneously. MOT performance was interfered by both the visual and the auditory task. However, the interference of visual task and auditory task with MOT occurred at different levels. The previous studies demonstrated that the MOT task and visual task occupy visual attention resources. Although the MOT task and auditory task don't occupy visual attention resources, they share more central attention resources(such as executive function). Multiple Identity Tracking(MIT) is a variant of MOT, in which each object carries a unique identity. The previous studies also included both visual or auditory task and MIT task, but those studies have not examined how the visual/auditory task affects the MIT task when the two tasks shared the same properties. The current study included 3 experiments and aimed to investigate the influence of a visual or auditory task on either MOT or MIT task. The first two experiments manipulated participants' eye movement to compare the different effects of visual task and auditory task on MOT performance. The result of experiment 1A showed that the auditory task interfered more with MOT than did the visual task when eyes were fixated at the center of the screen. However, the auditory task yielded less interference with MOT when there was no eye movement control in experiment 1B. The results indicated that the tracking processes in MOT not only occupy visual attention resources, but also occupy central attention resources(such as executive function). The experiment 2 applied visual and auditory digit judgment task and MIT task. When the object identity(marked by a digit) was identical to the number in the digit judgment, the MIT performance was facilitated by both the visual and auditory task. It is possible that the process of digit activated the target identities(the same digits) that was stored in visual working memory during tracking.
Read moreReducing MHT computational requirements through use of cheap JPDA methods
Hypothesis formation is a major computational burden for any multiple hypotheses tracking (MHT) method. In particular, a track-oriented MHT method defines compatible tracks to be tracks not sharing common observations and then re-forms hypotheses from compatible tracks after each new scan of data is received. The Cheap Joint Probabilistic Data Association (CJPDA) method provides an efficient means for computing approximate hypothesis probabilities. This paper presents a method of extending CJPDA calculations in order to eliminate low probability track branches in a track-oriented MHT method. The method is tested using IRST data. This approach reduces the number of tracks in a cluster and the resultant computations required for hypothesis formation. It is also suggested that the use of CJPDA methods can reduce assignment matrix sizes and resultant computations for the hypothesis-oriented (Reid’s algorithm) MHT implementation.
Read moreIntegrating the Reconstructed Scattering Center Feature Maps With Deep CNN Feature Maps for Automatic SAR Target Recognition
Automatic target recognition has been one of the hottest research in synthetic aperture radar (SAR) data processing. Noticing that popular recognition methods cannot utilize multiple features of SAR complex data, a method fused scattering center feature and deep convolutional neural network (CNN) feature is proposed in this letter. This method contains three key parts, namely, scattering center extraction and reconstruction block, CNN feature extraction block, and final feature fusion and classification block. In this process, the scattering center feature and CNN feature are fused at the level of feature maps, which retain the space information of 2-D feature maps. What is more, the proposed half end-to-end strategy realizes the automatic update of weighting parameters in feature extraction network and subnetwork, which promotes a better recognition efficiency. Experimental results on measured SAR data show that the proposed method can achieve better accuracy than other single feature-based methods and feature fusion methods.
Read moreNear Duplicate Image Retrieval using Multilevel Local and Global Convolutional Neural Network Features
In this work, we present an approach based on multilevel local as well as global Convolutional Neural Network (CNN) feature matching to retrieve near duplicate images. CNN features are suitable for visual matching. The CNN features of entire image may not give accuracy in retrieval due to various image editing/capturing operations. Our retrieval task focuses on matching image pairs based on local and global levels. In local matching, an image is segmented into fixed size blocks followed by extracting patches by considering neighboring regions at different levels. Matching local image patches at different levels provides robustness to our retrieval model. In local patch extraction, we select blocks containing SURF feature points instead of selecting all blocks. CNN features are extracted and stored for each image patch and then followed by extraction of global CNN features. Finally, similarity between image pairs is computed by considering all extracted CNN features. Our similarity function is based on correlation and number of blocks found in matching. We implemented our proposed approach on benchmarking Holiday dataset. Retrieval results show remarkable improvement in mean average precision (mAP) on the dataset.
Read moreA Novel Global Localization Approach Based on Structural Unit Encoding and Multiple Hypothesis Tracking
In this paper, we present a novel 2-D laser-based global localization approach for mobile robots, which is composed of geometrical relationship construction, a new structural unit encoding scheme (SUES), and an extended multiple hypothesis tracking (MHT) algorithm. Different from existing methods, we construct a 3-D "directional endpoint" feature encapsulating both the endpoint and the direction of a line segment; on this basis, a novel and efficient online structural unit encoding scheme (SUES) is proposed to describe the geometric relationship between the two directional endpoint features with some robustness to dynamic disturbances. Note that SUES presented in this paper is different from the bag-of-words scheme in two aspects: 1) SUES quantizes the geometrical relationship without offline training for vocabulary and 2) SUES is independent of the quality of the vocabulary. By factoring the global localization problem into a discrete pose estimation problem, the MHT is extended on the basis of SUES and odometry information to gradually restore the global robot pose. Different from the classical MHT framework, the extended MHT takes the independent observation results as the a priori, while the likelihood term is composed of consecutive candidate poses and the odometry information. Evaluations are carried out by using both publicly available data sets and self-recorded data sets. Comparative experimental results with respect to the adaptive Monte Carlo localization are presented to show the superior performance of the proposed approach in terms of success ratio and efficiency.
Read more