• Home
  • Search
  • A Novel Augmentative Backward Reward Function with Deep Reinforcement Learning for Autonomous UAV Navigation
  • Cite Icon7
  • https://doi.org/10.1080/08839514.2022.2084473Copy DOI Icon

A Novel Augmentative Backward Reward Function with Deep Reinforcement Learning for Autonomous UAV Navigation

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

ABSTRACT The autonomous UAV (unmanned aerial vehicle) navigation has recently gained an increasing interest from both academic and industrial sectors due to its potential uses in various fields and especially, the need for social distancing during the pandemic. Many works have adopted a deep reinforcement learning (RL) method with experience replay called deep deterministic policy gradient (DDPG) to control the motion of UAV, and gain high accuracy results in static and simplified environments. However, they are still far from being ready for real world adoption in that the UAVs have to operate under complex and dynamic conditions. We also found that using only DDPG makes the learning process prone to oscillation and is inefficient for tasks having high dimensional action-state spaces. Furthermore, the goal reward mechanism in traditional reward functions brings a bias to the state, which resembles the one at the goal area and leads to erroneous action selection. To get closer to being ready for real world adoption, we proposed a novel method that enables UAVs to be capable of handling motion control in realistic environments. The first component of our proposed method is point cloud data (PCD) simplification with truncated icosahedron structure which converts enormous PCD into a few essential data points. In the second component of our method, we replace the traditional goal reward mechanism with a new mechanism called Augmentative Backward Reward (ABR) function to dispense the goal reward to transitions proportionately to its participation. By integrating simplified PCD and ABR, we achieved significantly better results when compared with using only the-state-of-the-art, TD3. In addition, we tested the proposed method with another navigation task, BipedalWalkerHardcore, a testbed for RL, and the result is still better and steadier than of TD3. These results indicate that the proposed method is robust.

Similar Papers
  • Conference Article
  • Citations3

Autonomous UAV with Learned Trajectory Generation and Control

  • Oct 01, 2019
  • Yilan Li +4
  • Research Article
  • Citations47

Double Critic Deep Reinforcement Learning for Mapless 3D Navigation of Unmanned Aerial Vehicles

  • Jan 31, 2022
  • Journal of Intelligent & Robotic Systems
  • Ricardo Bedin Grando +4
  • Book Chapter

DeepFusion-NavNet: A Deep Learning Framework Combining Semantic Segmentation and Reinforcement Learning for Robust Autonomous UAV Navigation

  • Mar 19, 2026
  • Sukumar Rajendran +1
  • Book Chapter
  • Citations3

Responsive Regulation of Dynamic UAV Communication Networks Based on Deep Reinforcement Learning

  • Jan 01, 2022
  • Ran Zhang +4
  • Conference Article
  • Citations6

SREC: Proactive Self-Remedy of Energy-Constrained UAV-Based Networks via Deep Reinforcement Learning

  • Dec 01, 2020
  • Ran Zhang +2
  • PDF
  • Research Article
  • Citations13

Task Offloading Strategy for Unmanned Aerial Vehicle Power Inspection Based on Deep Reinforcement Learning

  • Mar 24, 2024
  • Sensors (Basel, Switzerland)
  • Wei Zhuang +2
  • Conference Article
  • Citations114

Autonomous navigation of UAV in large-scale unknown complex environment with deep reinforcement learning

  • Nov 01, 2017
  • Chao Wang +3
  • Conference Article
  • Citations36

A UAV Path Planning Method Based on Deep Reinforcement Learning

  • Jul 05, 2020
  • Yibing Li +4
  • Research Article
  • Citations14

MASAC-based confrontation game method of UAV clusters

  • Dec 01, 2022
  • SCIENTIA SINICA Informationis
  • 健 薛 +6
  • Research Article
  • Citations262

Study on deep reinforcement learning techniques for building energy consumption forecasting

  • Dec 03, 2019
  • Energy and Buildings
  • Tao Liu +4
  • Conference Article
  • Citations10

The Algorithm of UAV Automatic Landing System Using Computer Vision

  • May 01, 2020
  • Kostiantyn Dergachov +2
  • PDF
  • Research Article
  • Citations22

A Review of Mobile Robot Path Planning Based on Deep Reinforcement Learning Algorithm

  • Dec 01, 2021
  • Journal of Physics: Conference Series
  • Yanwei Zhao +2
  • Conference Article
  • Citations16

Joint Trajectory and Power Optimization for Energy Efficient UAV Communication Using Deep Reinforcement Learning

  • May 10, 2021
  • Yuling Cui +3
  • PDF
  • Research Article
  • Citations23

Efficient Deep Reinforcement Learning for Optimal Path Planning

  • Nov 07, 2022
  • Electronics
  • Jing Ren +2
  • Research Article
  • Citations6

Break through the limits of learning by machines

  • Sep 20, 2016
  • Chinese Science Bulletin
  • Zhongzhi Shi
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.