- Research Article
56
- 10.1016/j.neucom.2024.127760
A comprehensive overview of core modules in visual SLAM framework
- Apr 25, 2024
- Neurocomputing
- Dupeng Cai + 5 more +5
A comprehensive overview of core modules in visual SLAM framework
Visual simultaneous localization and mapping (VSLAM) is an essential technique used in areas such as robotics and augmented reality for pose estimation and 3D mapping. Research on VSLAM using both monocular and stereo cameras has grown significantly over the last two decades. There is, therefore, a need for emphasis on a comprehensive review of the evolving architecture of such algorithms in the literature. Although VSLAM algorithm pipelines share similar mathematical backbones, their implementations are individualized and the ad hoc nature of the interfacing between different modules of VSLAM pipelines complicates code reuseability and maintenance. This paper presents a software model for core components of VSLAM implementations and interfaces that govern data flow between them while also attempting to preserve the elements that offer performance improvements over the evolution of VSLAM architectures. The framework presented in this paper employs principles from model-driven engineering (MDE), which are used extensively in the development of large and complicated software systems. The presented VSLAM framework will assist researchers in improving the performance of individual modules of VSLAM while not having to spend time on system integration of those modules into VSLAM pipelines.
Loading PDF
A comprehensive overview of core modules in visual SLAM framework
A comprehensive overview of core modules in visual SLAM framework
3D Visual SLAM Based on Multiple Iterative Closest Point
With the development of novel RGB-D visual sensors, data association has been a basic problem in 3D Visual Simultaneous Localization and Mapping (VSLAM). To solve the problem, a VSLAM algorithm based on Multiple Iterative Closest Point (MICP) is presented. By using both RGB and depth information obtained from RGB-D camera, 3D models of indoor environment can be reconstructed, which provide extensive knowledge for mobile robots to accomplish tasks such as VSLAM and Human-Robot Interaction. Due to the limited views of RGB-D camera, additional information about the camera pose is needed. In this paper, the motion of the RGB-D camera is estimated by a motion capture system after a calibration process. Based on the estimated pose, the MICP algorithm is used to improve the alignment. A Kinect mobile robot which is running Robot Operating System and the motion capture system has been used for experiments. Experiment results show that not only the proposed VSLAM algorithm achieved good accuracy and reliability, but also the 3D map can be generated in real time.
Read moreA Robust RGB-D SLAM System for Indoor Environments With Reflective Ground
Visual Simultaneous Localization and Mapping (VSLAM) is a critical technology for intelligent mobile robots, enabling simultaneous environmental mapping and self-localization. However, reflective ground in indoor environments presents significant challenges to VSLAM systems, as it introduces noisy virtual landmarks that degrade both map accuracy and localization precision. To address these challenges, we propose a novel RGB-D VSLAM system specifically designed for these environments. Our approach begins with a carefully designed point-line-plane-based VSLAM framework, which mitigates the limitations caused by sparse point features in indoor settings. Then we analyze how reflective ground degrades VSLAM performance and propose a reflective ground feature removal strategy that integrates semantic and geometric information. To enhance robustness in highly reflective environments, we incorporate a monocular depth estimation network as a complementary module. Extensive experiments on public datasets demonstrate that our framework achieves state-of-the-art performance. And evaluations on an author-collected dataset highlight the system’s superior mapping and localization capabilities in reflective indoor environments. Furthermore, ablation studies validate the effectiveness of each proposed component, and latency tests confirm the system’s suitability for real-time applications.
Read moreEvent-Based Visual Simultaneous Localization and Mapping (EVSLAM) Techniques: State of the Art and Future Directions
Recent advances in event-based cameras have led to significant developments in robotics, particularly in visual simultaneous localization and mapping (VSLAM) applications. This technique enables real-time camera motion estimation and simultaneous environment mapping using visual sensors on mobile platforms. Event cameras offer several distinct advantages over frame-based cameras, including a high dynamic range, high temporal resolution, low power consumption, and low latency. These attributes make event cameras highly suitable for addressing performance issues in challenging scenarios such as high-speed motion and environments with high-range illumination. This review paper delves into event-based VSLAM (EVSLAM) algorithms, leveraging the advantages inherent in event streams for localization and mapping endeavors. The exposition commences by explaining the operational principles of event cameras, providing insights into the diverse event representations applied in event data preprocessing. A crucial facet of this survey is the systematic categorization of EVSLAM research into three key parts: event preprocessing, event tracking, and sensor fusion algorithms in EVSLAM. Each category undergoes meticulous examination, offering practical insights and guidance for comprehending each approach. Moreover, we thoroughly assess state-of-the-art (SOTA) methods, emphasizing conducting the evaluation on a specific dataset for enhanced comparability. This evaluation sheds light on current challenges and outlines promising avenues for future research, emphasizing the persisting obstacles and potential advancements in this dynamically evolving domain.
Read moreIntegration of BIM and ar with VSLAM to assist in construction site inspection
Building Information Modeling (BIM) has been widely adopted for construction inspections due to its ability to integrate multiple data sources. Engineers use BIM to identify and review site issues, yet inspection systems face several challenges. Firstly, positioning inspection areas on a construction site using BIM with Augmented Reality (AR) requires complex model manipulation. Additionally, signal or Internet connectivity issues may limit positioning technologies. Secondly, human error or interference is common in traditional inspection processes due to their complexity. To overcome these barriers, this research applied BIM and AR with Visual Simultaneous Localization and Mapping (VSLAM) to help inspectors quickly and effectively record construction defects as photographs with notes and their locations. An efficient approach is proposed to integrate BIM and AR with VSLAM, and a prototype is developed to validate and demonstrate how the proposed system can assist a site inspector in performing quality management, even offline. The system uses a two-phase indoor positioning method: initial localization via visual markers and real-time tracking with VSLAM, enabling precise defect tracking and efficient model adjustments. While significantly improving inspection accuracy and efficiency, its performance is affected by environmental factors like lighting and marker placement, providing insights for future refinement.
Read moreSemantic SLAM with Geometric Constraints for Object Tracking in Dynamic Environments
Visual simultaneous localization and mapping (VSLAM) is extensively utilized in the field of unmanned driving. Classical VSLAM typically postulates that the surroundings are stationary, and dynamic objects can seriously affect the accuracy of the system. However, detecting and tracking dynamic objects is beneficial for improving the autonomy and intelligence of robotics. In order to solve the problem of VSLAM distinguishing dynamic objects in dynamic scenes, a semantic VSLAM system based on the VDO-SLAM framework is proposed. The approach integrates semantic segmentation into the visual odometry and combines scene flow and geometric constraints to effectively reduce the negative impact of dynamic objects on system performance. Specifically, a transformer-based neural network architecture for optical flow estimation was employed, which effectively enhances the segmentation precision of foreground and background regions. And a geometric constraint was introduced to classify dynamic objects, which can reduction of erroneous segmentation of dynamic objects with effect. The data analysis results on the KITTI dataset demonstrate that our proposed algorithms can effectively enhance precision and stability of VSLAM system, reduce the tracking error of VSLAM in dynamic environments, and demonstrate good potential and application value.
Read moreAn Observer Design for Visual Simultaneous Localisation and Mapping with Output Equivariance
An Observer Design for Visual Simultaneous Localisation and Mapping with Output Equivariance
Eliminating short-term dynamic elements for robust visual simultaneous localization and mapping using a coarse-to-fine strategy
Visual simultaneous localization and mapping (VSLAM) is one of the foremost principal technologies for intelligent robots to implement environment perception. Many research works have focused on proposing comprehensive and integrated systems based on the static environment assumption. However, the elements whose motion status changes frequently, namely short-term dynamic elements, can significantly affect the system performance. Therefore, it is extremely momentous to cope with short-term dynamic elements to make the VSLAM system more adaptable to dynamic scenes. This paper proposes a coarse-to-fine elimination strategy for short-term dynamic elements based on motion status check (MSC) and feature points update (FPU). First, an object detection module is designed to obtain semantic information and screen out the potential short-term dynamic elements. And then an MSC module is proposed to judge the true status of these elements and thus ultimately determine whether to eliminate them. In addition, an FPU module is introduced to update the extracted feature points according to calculating the dynamic region factor to improve the robustness of VSLAM system. Quantitative and qualitative experiments on two challenging public datasets are performed. The results demonstrate that our method effectively eliminates the influence of short-term dynamic elements and outperforms other state-of-the-art methods.
Read moreStereo camera visual SLAM with hierarchical masking and motion-state classification at outdoor construction sites containing large dynamic objects
At modern construction sites, utilizing GNSS (Global Navigation Satellite System) to measure the real-time location and orientation (i.e. pose) of construction machines and navigate them is very common. However, GNSS is not always available. Replacing GNSS with on-board cameras and visual simultaneous localization and mapping (visual SLAM) to navigate the machines is a cost-effective solution. Nevertheless, at construction sites, multiple construction machines will usually work together and side-by-side, causing large dynamic occlusions in the cameras' view. Standard visual SLAM cannot handle large dynamic occlusions well. In this work, we propose a motion segmentation method to efficiently extract static parts from crowded dynamic scenes to enable robust tracking of camera ego-motion. Our method utilizes semantic information combined with object-level geometric constraints to quickly detect the static parts of the scene. Then, we perform a two-step coarse-to-fine ego-motion tracking with reference to the static parts. This leads to a novel dynamic visual SLAM formation. We test our proposals through a real implementation based on ORB-SLAM2, and datasets we collected from real construction sites. The results show that when standard visual SLAM fails, our method can still retain accurate camera ego-motion tracking in real-time. Comparing to state-of-the-art dynamic visual SLAM methods, ours shows outstanding efficiency and competitive result trajectory accuracy.
Read moreMOD-SLAM:Visual SLAM with Moving Object Detection in Dynamic Environments
In recent years, the signihcant progress has been made in visual simultaneous localization and mapping(VSLAM). Many present geometric VSLAM systems rely on static and bright environment assumptions, which is not friendly to the generalization of VSLAM in the real world including a large number of challenging scenes. To cope with these challenges, a real-time and robust visual-inertial SLAM system was proposed, which integrates a neural network for moving object detection(MOD) and greatly reduces the negative influence of dynamic objects. We has performed an ablation study to validate the effectiveness and necessity of our proposal. In addition, empirical evaluations on typical datasets, as well as in some usual dynamic environments, show that our novel framework can favorably solve the tracking loss, yield pure point cloud and improve the accuracy of VSLAM.
Read moreAsynchronous Multicamera SLAM Using Sparse Gaussian Process Regression
Visual Simultaneous Localization and Mapping (VSLAM) is a key technique that enables autonomous systems to localize themselves and incrementally build a map of an unknown environment using only visual information. Despite its importance, conventional VSLAM systems frequently experience drift and tracking failures in complex environments, which limits their overall effectiveness. Multi-camera systems have enhanced VSLAM accuracy by incorporating diverse viewpoints. However, they generally rely on synchronous data capture, restricting their applicability in multi-sensor setups where asynchronous data acquisition is necessary. To address this limitation, a continuous-time Asynchronous Multi-Camera SLAM (AMC-SLAM) framework is proposed, utilizing sparse Gaussian process regression. The method integrates outlier removal, continuous-time trajectory optimization, and multi-view loop closing to achieve robust pose estimation. Combining Gaussian process interpolation with bundle adjustment strengthens inter-camera data correlation and reduces the number of state variables, while the derived analytical Jacobians enhance optimization efficiency. Additionally, online estimation of multi-camera extrinsic parameters is incorporated to improve the system’s generality. Experimental results on the AMV-Bench dataset [1] demonstrate an absolute translation error of less than 0.5% over a 10 km trajectory, indicating significant improvements in accuracy over existing stereo and multi-camera SLAM systems. The method also exhibits robustness and generalizability on the New College Dataset [2] and in real-world scenarios. These findings underscore the potential of AMC-SLAM for high-precision, robust applications in challenging environments.
Read moreDEG-SLAM: a dynamic visual RGB-D SLAM based on object detection and geometric constraints for degenerate motion
Current visual simultaneous localization and mapping (SLAM) systems have demonstrated commendable efficacy in static environments. However, the presence of dynamic objects in real-world settings frequently leads to system discrepancies, significantly impairing the accuracy and robustness of SLAM systems. Conventional visual SLAM approaches typically utilize epipolar constraints to mitigate the impact of outliers; nevertheless, they encounter limitations when confronted with a substantial number of dynamic or planar moving objects. To tackle these issues, this paper introduces a novel dynamic visual SLAM system, termed DEG-SLAM. Initially, the system employs the YOLOv5 object detection network to identify dynamic objects, subsequently relaying the semantic information to the tracking module. During the tracking phase, both semantic information and epipolar constraints are leveraged to filter out dynamic feature points. To address the challenges posed by the malfunctioning of epipolar constraints in degenerate scenes, DEG-SLAM incorporates a degenerate constraint mechanism aimed at further eliminating dynamic feature points. Furthermore, a reprojection constraint has been introduced to enhance the filtering of absent dynamic feature points that lie outside the detection boxes. Experimental findings reveal that DEG-SLAM significantly improves accuracy and robustness when compared to traditional ORB-SLAM3 in dynamic environments. The performance benefits of DEG-SLAM are particularly evident in degenerate scenarios, thereby affirming its practicality and reliability in complex settings.
Read moreReal-Time 3D Reconstruction on Construction Site Using Visual SLAM and UAV
3D reconstruction can be used as a platform to monitor the performance of\nactivities on construction site, such as construction progress monitoring,\nstructure inspection and post-disaster rescue. Comparing to other sensors, RGB\nimage has the advantages of low-cost, texture rich and easy to implement that\nhas been used as the primary method for 3D reconstruction in construction\nindustry. However, the image-based 3D reconstruction always requires extended\ntime to acquire and/or to process the image data, which limits its application\non time critical projects. Recent progress in Visual Simultaneous Localization\nand Mapping (SLAM) make it possible to reconstruct a 3D map of construction\nsite in real-time. Integrated with Unmanned Aerial Vehicle (UAV), the obstacles\nareas that are inaccessible for the ground equipment can also be sensed.\nDespite these advantages of visual SLAM and UAV, until now, such technique has\nnot been fully investigated on construction site. Therefore, the objective of\nthis research is to present a pilot study of using visual SLAM and UAV for\nreal-time construction site reconstruction. The system architecture and the\nexperimental setup are introduced, and the preliminary results and the\npotential applications using Visual SLAM and UAV on construction site are\ndiscussed.\n
Read moreRETRACTED: A Synchronous Localization Method for Multiple Mobile Robots Based on Visual SLAM Algorithm in the Context of the Internet of Things
Following an investigation undertaken by the publisher, we have determined that this paper was accepted on the basis of a compromised peer review process. We hereby retract the paper. The corresponding author has been notified of the retraction. The retraction statement can be found here: https://doi.org/10.1520/JTE20269995. Aiming at the problem that traditional methods cannot meet the problem of synchronous positioning of multi-mobile robots in different motion states, a synchronous positioning method of multi-mobile robots based on visual simultaneous localization and mapping (SLAM) algorithm is proposed. Analyze the general model of SLAM problem, realize visual SLAM feature detection through scale invariant feature transform algorithm, and match image features based on active vision. A multi-mobile robot motion state estimation model is constructed, an octree model is used to construct a multi-robot synchronous localization map, and the camera pose tracking is realized by using the oriented fast accelerated segment test and rotated binary robust independent elementary features points that are not in the target bounding box. The three-dimensional points in the environment recovered by visual SLAM separate the background target and the moving target, and match the corresponding target in the previous frame through the feature to realize the synchronous positioning of the multi-mobile robot. The experimental results show that the method can realize the synchronous positioning of multi-robots under the state of multi-robot linear motion, curvilinear motion and mixed motion, and the positioning accuracy is high.
Read moreToward Accurate, Efficient, and Robust RGB-D Simultaneous Localization and Mapping in Challenging Environments
Visual Simultaneous Localization and Mapping (SLAM) is crucial to many applications such as self-driving vehicles and robot tasks. However, it is still challenging for existing visual SLAM approaches to achieve good performance in low-texture or illumination-changing scenes. In recent years, some researchers have turned to edge-based SLAM approaches to deal with the challenging scenes, which are more robust than feature-based and direct SLAM methods. Nevertheless, existing edge-based methods are computationally expensive and inferior than other visual SLAM systems in terms of accuracy. In this study, we propose EdgeSLAM, a novel RGB-D edge-based SLAM approach to deal with challenging scenarios that is efficient, accurate, and robust. EdgeSLAM is built on two innovative modules: efficient edge selection and adaptive robust motion estimation. The edge selection module can efficiently select a small set of edge pixels, which significantly improves the computational efficiency without sacrificing the accuracy. The motion estimation module improves the system's accuracy and robustness by adaptively handling outliers in motion estimation. Extensive experiments were conducted on TUM RGBD, ICL-NUIM and ETH3D datasets, and experimental results show that EdgeSLAM significantly outperforms five state-of-the-art (SOTA) methods in terms of efficiency, accuracy, and robustness, which achieves 29.17% accuracy improvements with a high processing speed of up to 120 FPS and a high positioning success rate of 97.06%.
Read more