- Research Article
56
- 10.1016/j.neucom.2024.127760
A comprehensive overview of core modules in visual SLAM framework
- Apr 25, 2024
- Neurocomputing
- Dupeng Cai + 5 more +5
A comprehensive overview of core modules in visual SLAM framework
Visual Simultaneous Localization and Mapping (VSLAM) is a key technique that enables autonomous systems to localize themselves and incrementally build a map of an unknown environment using only visual information. Despite its importance, conventional VSLAM systems frequently experience drift and tracking failures in complex environments, which limits their overall effectiveness. Multi-camera systems have enhanced VSLAM accuracy by incorporating diverse viewpoints. However, they generally rely on synchronous data capture, restricting their applicability in multi-sensor setups where asynchronous data acquisition is necessary. To address this limitation, a continuous-time Asynchronous Multi-Camera SLAM (AMC-SLAM) framework is proposed, utilizing sparse Gaussian process regression. The method integrates outlier removal, continuous-time trajectory optimization, and multi-view loop closing to achieve robust pose estimation. Combining Gaussian process interpolation with bundle adjustment strengthens inter-camera data correlation and reduces the number of state variables, while the derived analytical Jacobians enhance optimization efficiency. Additionally, online estimation of multi-camera extrinsic parameters is incorporated to improve the system’s generality. Experimental results on the AMV-Bench dataset [1] demonstrate an absolute translation error of less than 0.5% over a 10 km trajectory, indicating significant improvements in accuracy over existing stereo and multi-camera SLAM systems. The method also exhibits robustness and generalizability on the New College Dataset [2] and in real-world scenarios. These findings underscore the potential of AMC-SLAM for high-precision, robust applications in challenging environments.
A comprehensive overview of core modules in visual SLAM framework
A comprehensive overview of core modules in visual SLAM framework
Modelling Software Architecture for Visual Simultaneous Localization and Mapping
Visual simultaneous localization and mapping (VSLAM) is an essential technique used in areas such as robotics and augmented reality for pose estimation and 3D mapping. Research on VSLAM using both monocular and stereo cameras has grown significantly over the last two decades. There is, therefore, a need for emphasis on a comprehensive review of the evolving architecture of such algorithms in the literature. Although VSLAM algorithm pipelines share similar mathematical backbones, their implementations are individualized and the ad hoc nature of the interfacing between different modules of VSLAM pipelines complicates code reuseability and maintenance. This paper presents a software model for core components of VSLAM implementations and interfaces that govern data flow between them while also attempting to preserve the elements that offer performance improvements over the evolution of VSLAM architectures. The framework presented in this paper employs principles from model-driven engineering (MDE), which are used extensively in the development of large and complicated software systems. The presented VSLAM framework will assist researchers in improving the performance of individual modules of VSLAM while not having to spend time on system integration of those modules into VSLAM pipelines.
Read moreVisual Hybrid SLAM: An Appearance-Based Approach to Loop Closure
This paper proposes an appearance-based method to detect loop closure in visual SLAM (Simultaneous Localization and Mapping). To solve this problem, we make use of omnidirectional images and the internal odometry captured by a robot in a real indoor environment. We build an appearance-based model and, subsequently, two maps of the environment are constructed, one metric and other topological with relationships between them. These relationships are updated in each step of our hybrid approach. The topological map is a graph built from the appearance information in the scenes. A new node is added when the new visual information is different enough from the previous information. At the same time, we check a possible topological loop closure with previous nodes. On the other hand we estimate the metric position of the new pose using a Monte-Carlo approach with the aim of building a metric map. The experimental results demonstrate the reasonable performance of our method.KeywordsAppearance-base descriptorOmnidirectional ImagesMonte-Carlo SLAMHybrid Metric-Topological mappingLoop closure detection
Read moreIntegration of BIM and ar with VSLAM to assist in construction site inspection
Building Information Modeling (BIM) has been widely adopted for construction inspections due to its ability to integrate multiple data sources. Engineers use BIM to identify and review site issues, yet inspection systems face several challenges. Firstly, positioning inspection areas on a construction site using BIM with Augmented Reality (AR) requires complex model manipulation. Additionally, signal or Internet connectivity issues may limit positioning technologies. Secondly, human error or interference is common in traditional inspection processes due to their complexity. To overcome these barriers, this research applied BIM and AR with Visual Simultaneous Localization and Mapping (VSLAM) to help inspectors quickly and effectively record construction defects as photographs with notes and their locations. An efficient approach is proposed to integrate BIM and AR with VSLAM, and a prototype is developed to validate and demonstrate how the proposed system can assist a site inspector in performing quality management, even offline. The system uses a two-phase indoor positioning method: initial localization via visual markers and real-time tracking with VSLAM, enabling precise defect tracking and efficient model adjustments. While significantly improving inspection accuracy and efficiency, its performance is affected by environmental factors like lighting and marker placement, providing insights for future refinement.
Read moreSemantic SLAM with Geometric Constraints for Object Tracking in Dynamic Environments
Visual simultaneous localization and mapping (VSLAM) is extensively utilized in the field of unmanned driving. Classical VSLAM typically postulates that the surroundings are stationary, and dynamic objects can seriously affect the accuracy of the system. However, detecting and tracking dynamic objects is beneficial for improving the autonomy and intelligence of robotics. In order to solve the problem of VSLAM distinguishing dynamic objects in dynamic scenes, a semantic VSLAM system based on the VDO-SLAM framework is proposed. The approach integrates semantic segmentation into the visual odometry and combines scene flow and geometric constraints to effectively reduce the negative impact of dynamic objects on system performance. Specifically, a transformer-based neural network architecture for optical flow estimation was employed, which effectively enhances the segmentation precision of foreground and background regions. And a geometric constraint was introduced to classify dynamic objects, which can reduction of erroneous segmentation of dynamic objects with effect. The data analysis results on the KITTI dataset demonstrate that our proposed algorithms can effectively enhance precision and stability of VSLAM system, reduce the tracking error of VSLAM in dynamic environments, and demonstrate good potential and application value.
Read moreA Robust RGB-D SLAM System for Indoor Environments With Reflective Ground
Visual Simultaneous Localization and Mapping (VSLAM) is a critical technology for intelligent mobile robots, enabling simultaneous environmental mapping and self-localization. However, reflective ground in indoor environments presents significant challenges to VSLAM systems, as it introduces noisy virtual landmarks that degrade both map accuracy and localization precision. To address these challenges, we propose a novel RGB-D VSLAM system specifically designed for these environments. Our approach begins with a carefully designed point-line-plane-based VSLAM framework, which mitigates the limitations caused by sparse point features in indoor settings. Then we analyze how reflective ground degrades VSLAM performance and propose a reflective ground feature removal strategy that integrates semantic and geometric information. To enhance robustness in highly reflective environments, we incorporate a monocular depth estimation network as a complementary module. Extensive experiments on public datasets demonstrate that our framework achieves state-of-the-art performance. And evaluations on an author-collected dataset highlight the system’s superior mapping and localization capabilities in reflective indoor environments. Furthermore, ablation studies validate the effectiveness of each proposed component, and latency tests confirm the system’s suitability for real-time applications.
Read moreDMOT-SLAM: visual SLAM in dynamic environments with moving object tracking
Visual simultaneous localization and mapping (SLAM) in dynamic environments has received significant attention in recent years, and accurate segmentation of real dynamic objects is the key to enhancing the accuracy of pose estimation in such environments. In this study, we propose a visual SLAM approach based on ORB-SLAM3, namely dynamic multiple object tracking SLAM (DMOT-SLAM), which can accurately estimate the camera’s pose in dynamic environments while tracking the trajectories of moving objects. We introduce a spatial point correlation constraint and combine it with instance segmentation and epipolar constraint to identify dynamic objects. We integrate the proposed motion check method into DeepSort, an object tracking algorithm, to facilitate inter-frame tracking of dynamic objects. This integration not only enhances the stability of dynamic features detection but also enables the estimation of global motion trajectories for dynamic objects and the construction of object-level semi-dense semantic maps. We evaluate our approach on the public TUM, Bonn, and KITTI dataset, and the results show that our approach has a significant improvement over ORB-SLAM3 in dynamic scenes and performs better compared to other state-of-the-art SLAM approaches. Moreover, experiments in real-world scenarios further substantiate the effectiveness of our approach.
Read moreDistilled representation using patch-based local-to-global similarity strategy for visual place recognition
Distilled representation using patch-based local-to-global similarity strategy for visual place recognition
DEG-SLAM: a dynamic visual RGB-D SLAM based on object detection and geometric constraints for degenerate motion
Current visual simultaneous localization and mapping (SLAM) systems have demonstrated commendable efficacy in static environments. However, the presence of dynamic objects in real-world settings frequently leads to system discrepancies, significantly impairing the accuracy and robustness of SLAM systems. Conventional visual SLAM approaches typically utilize epipolar constraints to mitigate the impact of outliers; nevertheless, they encounter limitations when confronted with a substantial number of dynamic or planar moving objects. To tackle these issues, this paper introduces a novel dynamic visual SLAM system, termed DEG-SLAM. Initially, the system employs the YOLOv5 object detection network to identify dynamic objects, subsequently relaying the semantic information to the tracking module. During the tracking phase, both semantic information and epipolar constraints are leveraged to filter out dynamic feature points. To address the challenges posed by the malfunctioning of epipolar constraints in degenerate scenes, DEG-SLAM incorporates a degenerate constraint mechanism aimed at further eliminating dynamic feature points. Furthermore, a reprojection constraint has been introduced to enhance the filtering of absent dynamic feature points that lie outside the detection boxes. Experimental findings reveal that DEG-SLAM significantly improves accuracy and robustness when compared to traditional ORB-SLAM3 in dynamic environments. The performance benefits of DEG-SLAM are particularly evident in degenerate scenarios, thereby affirming its practicality and reliability in complex settings.
Read moreToward Accurate, Efficient, and Robust RGB-D Simultaneous Localization and Mapping in Challenging Environments
Visual Simultaneous Localization and Mapping (SLAM) is crucial to many applications such as self-driving vehicles and robot tasks. However, it is still challenging for existing visual SLAM approaches to achieve good performance in low-texture or illumination-changing scenes. In recent years, some researchers have turned to edge-based SLAM approaches to deal with the challenging scenes, which are more robust than feature-based and direct SLAM methods. Nevertheless, existing edge-based methods are computationally expensive and inferior than other visual SLAM systems in terms of accuracy. In this study, we propose EdgeSLAM, a novel RGB-D edge-based SLAM approach to deal with challenging scenarios that is efficient, accurate, and robust. EdgeSLAM is built on two innovative modules: efficient edge selection and adaptive robust motion estimation. The edge selection module can efficiently select a small set of edge pixels, which significantly improves the computational efficiency without sacrificing the accuracy. The motion estimation module improves the system's accuracy and robustness by adaptively handling outliers in motion estimation. Extensive experiments were conducted on TUM RGBD, ICL-NUIM and ETH3D datasets, and experimental results show that EdgeSLAM significantly outperforms five state-of-the-art (SOTA) methods in terms of efficiency, accuracy, and robustness, which achieves 29.17% accuracy improvements with a high processing speed of up to 120 FPS and a high positioning success rate of 97.06%.
Read moreRETRACTED: A Synchronous Localization Method for Multiple Mobile Robots Based on Visual SLAM Algorithm in the Context of the Internet of Things
Following an investigation undertaken by the publisher, we have determined that this paper was accepted on the basis of a compromised peer review process. We hereby retract the paper. The corresponding author has been notified of the retraction. The retraction statement can be found here: https://doi.org/10.1520/JTE20269995. Aiming at the problem that traditional methods cannot meet the problem of synchronous positioning of multi-mobile robots in different motion states, a synchronous positioning method of multi-mobile robots based on visual simultaneous localization and mapping (SLAM) algorithm is proposed. Analyze the general model of SLAM problem, realize visual SLAM feature detection through scale invariant feature transform algorithm, and match image features based on active vision. A multi-mobile robot motion state estimation model is constructed, an octree model is used to construct a multi-robot synchronous localization map, and the camera pose tracking is realized by using the oriented fast accelerated segment test and rotated binary robust independent elementary features points that are not in the target bounding box. The three-dimensional points in the environment recovered by visual SLAM separate the background target and the moving target, and match the corresponding target in the previous frame through the feature to realize the synchronous positioning of the multi-mobile robot. The experimental results show that the method can realize the synchronous positioning of multi-robots under the state of multi-robot linear motion, curvilinear motion and mixed motion, and the positioning accuracy is high.
Read moreReal-Time 3D Reconstruction on Construction Site Using Visual SLAM and UAV
3D reconstruction can be used as a platform to monitor the performance of\nactivities on construction site, such as construction progress monitoring,\nstructure inspection and post-disaster rescue. Comparing to other sensors, RGB\nimage has the advantages of low-cost, texture rich and easy to implement that\nhas been used as the primary method for 3D reconstruction in construction\nindustry. However, the image-based 3D reconstruction always requires extended\ntime to acquire and/or to process the image data, which limits its application\non time critical projects. Recent progress in Visual Simultaneous Localization\nand Mapping (SLAM) make it possible to reconstruct a 3D map of construction\nsite in real-time. Integrated with Unmanned Aerial Vehicle (UAV), the obstacles\nareas that are inaccessible for the ground equipment can also be sensed.\nDespite these advantages of visual SLAM and UAV, until now, such technique has\nnot been fully investigated on construction site. Therefore, the objective of\nthis research is to present a pilot study of using visual SLAM and UAV for\nreal-time construction site reconstruction. The system architecture and the\nexperimental setup are introduced, and the preliminary results and the\npotential applications using Visual SLAM and UAV on construction site are\ndiscussed.\n
Read moreEvent-Based Visual Simultaneous Localization and Mapping (EVSLAM) Techniques: State of the Art and Future Directions
Recent advances in event-based cameras have led to significant developments in robotics, particularly in visual simultaneous localization and mapping (VSLAM) applications. This technique enables real-time camera motion estimation and simultaneous environment mapping using visual sensors on mobile platforms. Event cameras offer several distinct advantages over frame-based cameras, including a high dynamic range, high temporal resolution, low power consumption, and low latency. These attributes make event cameras highly suitable for addressing performance issues in challenging scenarios such as high-speed motion and environments with high-range illumination. This review paper delves into event-based VSLAM (EVSLAM) algorithms, leveraging the advantages inherent in event streams for localization and mapping endeavors. The exposition commences by explaining the operational principles of event cameras, providing insights into the diverse event representations applied in event data preprocessing. A crucial facet of this survey is the systematic categorization of EVSLAM research into three key parts: event preprocessing, event tracking, and sensor fusion algorithms in EVSLAM. Each category undergoes meticulous examination, offering practical insights and guidance for comprehending each approach. Moreover, we thoroughly assess state-of-the-art (SOTA) methods, emphasizing conducting the evaluation on a specific dataset for enhanced comparability. This evaluation sheds light on current challenges and outlines promising avenues for future research, emphasizing the persisting obstacles and potential advancements in this dynamically evolving domain.
Read moreAn Observer Design for Visual Simultaneous Localisation and Mapping with Output Equivariance
An Observer Design for Visual Simultaneous Localisation and Mapping with Output Equivariance
Eliminating short-term dynamic elements for robust visual simultaneous localization and mapping using a coarse-to-fine strategy
Visual simultaneous localization and mapping (VSLAM) is one of the foremost principal technologies for intelligent robots to implement environment perception. Many research works have focused on proposing comprehensive and integrated systems based on the static environment assumption. However, the elements whose motion status changes frequently, namely short-term dynamic elements, can significantly affect the system performance. Therefore, it is extremely momentous to cope with short-term dynamic elements to make the VSLAM system more adaptable to dynamic scenes. This paper proposes a coarse-to-fine elimination strategy for short-term dynamic elements based on motion status check (MSC) and feature points update (FPU). First, an object detection module is designed to obtain semantic information and screen out the potential short-term dynamic elements. And then an MSC module is proposed to judge the true status of these elements and thus ultimately determine whether to eliminate them. In addition, an FPU module is introduced to update the extracted feature points according to calculating the dynamic region factor to improve the robustness of VSLAM system. Quantitative and qualitative experiments on two challenging public datasets are performed. The results demonstrate that our method effectively eliminates the influence of short-term dynamic elements and outperforms other state-of-the-art methods.
Read more