- Research Article
1
- 10.1016/j.neucom.2021.08.008
Self-supervised multi-body scene flow estimation
- Aug 12, 2021
- Neurocomputing
- Jihuang Dai + 2 more +2
Self-supervised multi-body scene flow estimation
In the general structure-from-motion (SFM) problem involving several moving objects in a scene, the essential first step is to segment moving objects independently. We attempt to deal with the problem of optical flow estimation and motion segmentation over a pair of images. We apply a mean field technique to determine optical flow and motion boundaries and present a deterministic algorithm. Since motion discontinuities represented by line process are embedded in the estimation of the optical flow, our algorithm provides accurate estimates of optical flow especially along motion boundaries and handles occlusion and multiple motions. We show that the proposed algorithm outperforms other well-known algorithms in terms of estimation accuracy and timing.
Self-supervised multi-body scene flow estimation
Self-supervised multi-body scene flow estimation
Self-Supervised Optical Flow Estimation by Projective Bootstrap
Dense optical flow estimation is complex and time consuming, with state-of-the-art methods relying either on large synthetic data sets or on pipelines requiring up to a few minutes per frame pair. In this paper, we address the problem of optical flow estimation in the automotive scenario in a self-supervised manner. We argue that optical flow can be cast as a geometrical warping between two successive video frames and devise a deep architecture to estimate such transformation in two stages. First, a dense pixel-level flow is computed with a projective bootstrap on rigid surfaces. We show how such global transformation can be approximated with a homography and extend spatial transformer layers so that they can be employed to compute the flow field implied by such transformation. Subsequently, we refine the prediction by feeding a second, deeper network that accounts for moving objects. A final reconstruction loss compares the warping of frame X <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">t</sub> with the subsequent frame X <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">t+1</sub> and guides both estimates. The model has the speed advantages of end-to-end deep architectures while achieving competitive performances, both outperforming recent unsupervised methods and showing good generalization capabilities on new automotive data sets.
Read moreTotal least squares estimation of stereo optical flow
We propose a new method for disparity assisted stereo optical flow estimation. This method is based on the linearization of the round-about compatibility constraint, which converts a stereo optical flow estimation problem to a single channel optical flow estimation problem. An over-determined system of optical flow equations can then be constructed for estimating the flow fields. The total least squares, instead of the traditional least squares is used to solve the estimation problem. We also investigate the extension of the locally constant flow model across the time domain. The experiments presented demonstrate that the proposed technique performs very well.
Read moreRecursive Optical Flow Estimation—Adaptive Filtering Approach
Recursive Optical Flow Estimation—Adaptive Filtering Approach
UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos
In this paper, we propose UnOS, an unified system for unsupervised optical flow and stereo depth estimation using convolutional neural network (CNN) by taking advantages of their inherent geometrical consistency based on the rigid-scene assumption. UnOS significantly outperforms other state-of-the-art (SOTA) unsupervised approaches that treated the two tasks independently. Specifically, given two consecutive stereo image pairs from a video, UnOS estimates per-pixel stereo depth images, camera ego-motion and optical flow with three parallel CNNs. Based on these quantities, UnOS computes rigid optical flow and compares it against the optical flow estimated from the FlowNet, yielding pixels satisfying the rigid-scene assumption. Then, we encourage geometrical consistency between the two estimated flows within rigid regions, from which we derive a rigid-aware direct visual odometry (RDVO) module. We also propose rigid and occlusion-aware flow-consistency losses for the learning of UnOS. We evaluated our results on the popular KITTI dataset over 4 related tasks, \ie stereo depth, optical flow, visual odometry and motion segmentation.
Read moreOptical flow estimation based on local-global modeling and visual similarity guidance
Optical flow estimation is a research core of computer vision. In recent years, optical flow estimation methods based on convolutional neural networks (CNNs) have achieved great success. However, most of the existing CNN-based methods are incapable of modeling long-distance dependencies successfully due to the limited receptive fields of convolutions. This may lead to poor performance of optical flow estimation in regions of large displacements and local ambiguities. Furthermore, the problem of edge-blurring remains an open challenge for CNN-based optical flow methods because the general interpolation operation tends to amplify the optical flow errors during the upsampling process. To deal with the abovementioned problems, this paper proposes a novel optical flow estimation method based on local-global modeling and visual similarity guidance. First, an efficient and simple self-attention module is adopted to improve the local-global modeling ability of the network and to extract the more representative features, which decreases the number of optical flow errors caused by large displacements. Second, based on the assumption that the visual features of the objects are more similar, the motion is more similar, and a visual similarity-guided optical flow upsampling network is constructed to guide the upsampling process. By transforming the visual similarity of features into motion similarity, the proposed upsampling scheme improves the accuracy of optical flow estimation in regions of motion boundaries. Finally, we run the proposed method on MPI-Sintel and KITTI test datasets to conduct a comprehensive comparison with some state-of-the-art methods. The experimental results indicate that the proposed method achieves top performance among the comparable methods and, in particular, gains significant developments in regions of large displacements and motion boundaries.
Read moreOptical flow estimation in ultrasound images using a sparse representation
This paper introduces a 2D optical flow estimation method for cardiac ultrasound imaging based on a sparse representation. The optical flow problem is regularized using a classical gradient-based smoothness term combined with a sparsity inducing regularization that uses a learned cardiac flow dictionary. A particular emphasis is put on the influence of the spatial and sparse regularizations on the optical flow estimation problem. A comparison with state-of-the-art methods using realistic simulations shows the competitiveness of the proposed method for cardiac motion estimation in ultrasound images.
Read moreOptical Flow Estimation: An Error Analysis of Gradient-Based Methods with Local Optimization
Multiple views of a scene can provide important information about the structure and dynamic behavior of three-dimensional objects. Many of the methods that recover this information require the determination of optical flow-the velocity, on the image, of visible points on object surfaces. An important class of techniques for estimating optical flow depend on the relationship between the gradients of image brightness. While gradient-based methods have been widely studied, little attention has been paid to accuracy and reliability of the approach. Gradient-based methods are sensitive to conditions commonly encountered in real imagery. Highly textured surfaces, large areas of constant brightness, motion boundaries, and depth discontinuities can all be troublesome for gradient-based methods. Fortunately, these problematic areas are usually localized can be identified in the image. In this paper we examine the sources of errors for gradient-based techniques that locally solve for optical flow. These methods assume that optical flow is constant in a small neighborhood. The consequence of violating in this assumption is examined. The causes of measurement errors and the determinants of the conditioning of the solution system are also considered. By understanding how errors arise, we are able to define the inherent limitations of the technique, obtain estimates of the accuracy of computed values, enhance the performance of the technique, and demonstrate the informative value of some types of error.
Read moreFeature-Level Collaboration: Joint Unsupervised Learning of Optical Flow, Stereo Depth and Camera Motion
Precise estimation of optical flow, stereo depth and camera motion are important for the real-world 3D scene understanding and visual perception. Since the three tasks are tightly coupled with the inherent 3D geometric constraints, current studies have demonstrated that the three tasks can be improved through jointly optimizing geometric loss functions of several individual networks. In this paper, we show that effective feature-level collaboration of the networks for the three respective tasks could achieve much greater performance improvement for all three tasks than only loss-level joint optimization. Specifically, we propose a single network to combine and improve the three tasks. The network extracts the features of two consecutive stereo images, and simultaneously estimates optical flow, stereo depth and camera motion. The whole network mainly contains four parts: (I) a feature-sharing encoder to extract features of input images, which can enhance features' representation ability; (II) a pooled decoder to estimate both optical flow and stereo depth; (III) a camera pose estimation module which fuses optical flow and stereo depth information; (IV) a cost volume complement module to improve the performance of optical flow in static and occluded regions. Our method achieves state-of-the-art performance among the joint unsupervised methods, including optical flow and stereo depth estimation on KITTI 2012 and 2015 benchmarks, and camera motion estimation on KITTI VO dataset.
Read moreRobust discontinuity-preserving model for estimating optical flow
We address the problem of incremental optical flow estimation. Following Black et al., we design a cost function whose data and prior terms both involve robust M-estimators. The non-convex minimization is lead with an efficient multigrid algorithm which converges fast toward estimates of good quality while providing, at very low cost, crude estimates revealing the large discontinuity structures of the flow field.
Read moreGuided Filtering: Toward Edge-Preserving for Optical Flow
Despite progress made in the accuracy and robustness of optical flow in past years, the problem of over-segmentation and the blurring of image edge and motion boundary caused by the illumination change, complex texture, large displacement, and motion occlusion still remain. Recently, we developed a guided filtering scheme for flow field estimation, which is implemented as an add-on optimal operation during the coarse-to-fine optical flow computation. In this paper, we first review the research progress in optical flow computation and discuss limitations of the currently popular median filtering heuristic for a flow field optimization. We then introduce a general formulation of the guided filtering and provide the detailed illustration. Furthermore, we explore the potential of the guided filtering optimization for the flow field estimation under the coarse-to-fine computing scheme. Finally, we modify some typical and state-of-the-art optical flow methods by applying the proposed guided filtering operation to the baseline models, and test the performances of the basic and developed models through the Middlebury, MPI-Sintel, and KITTI data. The experimental results demonstrate that the guided filtering scheme is able to preserve the image edges and motion boundaries, and to improve the accuracy and robustness of optical flow estimation.
Read morePerceptual Loss for Convolutional Neural Network Based Optical Flow Estimation
Convolutional Neural Networks (CNNs) are successfully used in optical flow estimation as learned patch based descriptors. In this work, rather training feature descriptors via CNNs, an end-to-end fully convolutional network, is developed for solving optical flow from a pair of images. Motivated by the success in image transformation tasks, a perceptual loss function is used for training the network for optical flow estimation. We trained a deep convolutional auto-encoder of optical flow field to obtain the high-level representation of motion structures rather than image texture. The perceptual loss function is then defined by high-level features extracted from the pretrained encoder. Conventional variational refinement are not performed. Experiments show the network achieves competitive performance on the challenging MPI Sintel set and Flying Chairs set.
Read moreReal Time Dense Motion Estimation Using FPGA Based Omnidirectional Video Acquisition Device
OVAD (Fraś et al. in Vision based systems for UAV applications, Springer International Publishing, pp. 123–136 [1]; Kwiatkowski et al. in Recent advances in electrical engineering and computer science, pp. 58–61 [2]) is a device, which is purposed for multi directional video acquisition. Further development of the device includes hardware implementation of time-absorbing calculations, connected with video stream processing, with utilization of FPGA (Field Programmable Gate Array). The paper presents the results of research on the problem of real time dense optical flow estimation. Furthermore, the described results are connected to the problem of parallel processing and hardware acceleration using FPGA. Proposed solution may be found as useful in many applications, both civilian and military and proves utility of FPGA based solutions in real time video processing.
Read moreRicher Aggregated Features for Optical Flow Estimation with Edge-aware Refinement
Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attention to the design of flow decoder than the feature extraction. In this paper, we present a novel optical flow estimation network to enrich the feature representation of each pyramid level, with a hierarchical dilated architecture and a bottom-up aggregation scheme. In addition, inspired by edge guided classical methods, we bring the edge-aware idea into our approach and propose an edge-aware refinement (EAR) subnetwork to handle motion boundaries. Using the same decoding structure as PWC-Net, our network outperforms it by a large margin and leads all its derivatives both on KITTI-2012 and KITTI-2015. Further performance analysis proves the effectiveness of proposed ideas.
Read morePanoFlow: Learning 360° Optical Flow for Surrounding Temporal Understanding
Optical flow estimation is a basic task in self-driving and robotics systems, which enables to temporally interpret traffic scenes. Autonomous vehicles clearly benefit from the ultra-wide Field of View (FoV) offered by 360° panoramic sensors. However, due to the unique imaging process of panoramic cameras, models designed for pinhole images do not directly generalize satisfactorily to 360° panoramic images. In this paper, we put forward a novel network framework——PANO FLOW, to learn optical flow for panoramic images. To overcome the distortions introduced by equirectangular projection in panoramic transformation, we design a Flow Distortion Augmentation (FDA) method, which contains radial flow distortion (FDA-R) or equirectangular flow distortion (FDA-E). We further look into the definition and properties of cyclic optical flow for panoramic videos, and hereby propose a Cyclic Flow Estimation (CFE) method by leveraging the cyclicity of spherical images to infer 360° optical flow and converting large displacement to relatively small displacement. PanoFlow is applicable to any existing flow estimation method and benefits from the progress of narrow-FoV flow estimation. In addition, we create and release a synthetic panoramic dataset FlowScape based on CARLA to facilitate training and quantitative analysis. PanoFlow achieves state-of-the-art performance on the public OmniFlowNet and the fresh established FlowScape benchmarks. Our proposed approach reduces the End-Point-Error (EPE) on FlowScape by 27.3%. On OmniFlowNet, PanoFlow achieves an EPE of 3.17 pixels, a 55.5% error reduction from the best published result (7.12 pixels). We also qualitatively validate our method via an outdoor collection vehicle and a public real-world OmniPhotos dataset, indicating strong potential and robustness for real-world navigation applications. Code and dataset are publicly available at PanoFlow.
Read more