• Home
  • Search
  • Visual metric and semantic localization for UGV
  • https://doi.org/10.32657/10356/162442Copy DOI Icon

Visual metric and semantic localization for UGV

  • Jan 1, 2022
  • Handuo Zhang
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

During the continual transitions from lab research to real-world applications of vision-based algorithms, there are significant challenges, e.g. the robustness to adapt to complex environments and the high demands of multi-task, multi-modal learning models. More specifically on visual simultaneous localization and mapping (visual SLAM) for mobile robots, two notable limitations are: (i) the drift issue during pose estimation, especially in dynamic environments, makes the positioning system unstable; (ii) the additional loop detection takes many computational resources for image retrieval and global feature geometry checking. This thesis explores the approaches of visual localization on unmanned ground vehicles by solving those limitations. The conventional feature extraction and matching pipeline for vision-based tasks involve three steps: (i) local feature detection and description, (ii) feature matching, and (iii) outlier rejection. The first contribution of the thesis is adding a feature selection and anticipation stage to reduce tracking drift. We explore the noise model of image features to select a subset of all the observed image features with the best ”contribution” during data association and pose estimation across multiple frames. Conventional SLAM algorithms take a strong assumption of scene rigidity, which limits the application under challenging environments. The second part of the thesis addresses the tough issue in dynamic environments with moving objects. We presented GMC, namely the motion clustering approach, a lightweight dynamic object filtering method. It can distinguish moving objects from static landmarks. Based on the theory of motion coherence within a particular image area, GMC could segment dynamic objects in 3D space. We can provide an efficient and robust correspondence algorithm that can extract dynamic objects from a static background with the method. In this way, we propose a dynamic SLAM system that is real-time and free from expensive GPU processors. In contrast to GMC, the thesis’s third part turns to an object-aware learning-based model for more general dynamic scenarios. We use object detection and tracking as points, lines, planes, etc. We utilize semantic information and extract sparse image features simultaneously to keep track of dynamic objects. The static background and different dynamic objects are jointly optimized in a newly developed bundle adjustment sliding window. The estimated 3D bounding boxes can provide more robust camera tracking and better scene understanding, and better map merging. The fourth part of the thesis leverages the emerging feature learning framework. It proposes a unified self-supervised model called LGDNet to generate both global and local image feature descriptors end-to-end. Global feature descriptors embed the whole image into a compact representation, leading to easier scene comparison. On the other hand, local features focus more on the local region similarities of some salient parts for structure from motion. Our proposed method can directly extract features together with descriptors that encode both local maximum responses and global context information, avoiding duplicate calculations based on different feature extraction criteria.

Similar Papers
  • Peer Review Report

Decision letter: Visual and motor signatures of locomotion dynamically shape a population code for feature detection in Drosophila

  • Oct 03, 2022
  • Terufumi Fujiwara +1
  • Research Article
  • Citations64

False-positive reduction technique for detection of masses on digital mammograms: Global and local multiresolution texture analysis

  • Jun 01, 1997
  • Medical Physics
  • Datong Wei +6
  • Conference Article
  • Citations9

CNN classification based on global and local features

  • May 14, 2019
  • Yufeng Zheng +4
  • Conference Article
  • Citations6

Cascading global and local features for face recognition using support vector machines and local ternary patterns

  • Jul 01, 2017
  • Jia-Ching Jang Jian +5
  • PDF
  • Research Article
  • Citations18

Combining Local and Global Features Into a Siamese Network for Sentence Similarity

  • Jan 01, 2020
  • IEEE Access
  • Yulong Li +2
  • Conference Article
  • Citations1

Deep Single Image Enhancer

  • Sep 01, 2019
  • Mengchen Lin +2
  • Conference Article
  • Citations59

Hierarchical Ensemble of Global and Local Classifiers for Face Recognition

  • Jan 01, 2007
  • Yu Su +3
  • Research Article
  • Citations3

Toward Deeper Understanding of Children’s Writing: Pre-Service Teachers’ Attention to Local and Global Text Features at the Start and End of Writing-Focused Coursework

  • Aug 15, 2021
  • Literacy Research and Instruction
  • Lisa K Hawkins +4
  • Research Article
  • Citations85

An Efficient and Robust Framework for SAR Target Recognition by Hierarchically Fusing Global and Local Features.

  • Aug 03, 2018
  • IEEE Transactions on Image Processing
  • Baiyuan Ding +3
  • Book Chapter
  • Citations20

Gender Recognition Using Fusion of Local and Global Facial Features

  • Jan 01, 2013
  • Anwar M Mirza +5
  • Research Article
  • Citations8

Intelligent Fault Diagnosis Method for Shearer Rocker Gear Based on Swin Transformer and Multiscale Convolution Parallel Integration

  • Jan 01, 2025
  • IEEE Transactions on Instrumentation and Measurement
  • Xiaochun Sun +5
  • Research Article
  • Citations4

Face recognition fusing global and local features

  • Jan 01, 2006
  • Journal of Electronic Imaging
  • Wei-Wei Yu
  • Research Article
  • Citations4

Dual Backbone Multi-Attention Hierarchical Fusion and Feature Enhancement Network for Crowd Counting

  • May 01, 2025
  • IEEE Transactions on Consumer Electronics
  • Chunling Zheng +3
  • PDF
  • Research Article
  • Citations23

A Structure-Based B-cell Epitope Prediction Model Through Combing Local and Global Features

  • Jul 01, 2022
  • Frontiers in Immunology
  • Shuai Lu +4
  • Research Article
  • Citations29

BDNet: A BERT-based dual-path network for text-to-image cross-modal person re-identification

  • Apr 24, 2023
  • Pattern Recognition
  • Qiang Liu +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.