• Home
  • Search
  • Heuristic Function Negotiation for Markov Decision Process and Its Application in UAV Simulation
  • Cite Icon2
  • https://doi.org/10.1587/transinf.e97.d.89Copy DOI Icon

Heuristic Function Negotiation for Markov Decision Process and Its Application in UAV Simulation

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The traditional reinforcement learning (RL) methods can solve Markov Decision Processes (MDPs) online, but these learning methods cannot effectively use a priori knowledge to guide the learning process. The exploration of the optimal policy is time-consuming and does not employ the information about specific issues. To tackle the problem, this paper proposes heuristic function negotiation (HFN) as an online learning framework. The HFN framework extends MDPs and introduces heuristic functions. HFN changes the state-action dual layer structure of traditional RL to the triple layer structure, in which multiple heuristic functions can be set to meet the needs required to solve the problem. The HFN framework can use different algorithms to let the functions negotiate to determine the appropriate action, and adjust the impact of each function according to the rewards. The HFN framework introduces domain knowledge by setting heuristic functions and thus speeds up the problem solving of MDPs. Furthermore, user preferences can be reflected in the learning process, which improves the flexibility of RL. The experiments show that, by setting reasonable heuristic functions, the learning results of the HFN framework are more efficient than traditional RL. We also apply HFN to the air combat simulation of unmanned aerial vehicles (UAVs), which shows that different function settings lead to different combat behaviors.

Similar Papers
  • Research Article
  • Citations13

Reinforcement Learning for Clinical Applications.

  • Feb 08, 2023
  • Clinical Journal of the American Society of Nephrology
  • Kia Khezeli +5
  • Conference Article

QMDP: DASH Adaptation using Queueing Theory within a Markov Decision Process

  • Jan 09, 2021
  • Kevin Gatimu +1
  • Book Chapter
  • Citations17

Solving Markov Decision Processes via Simulation

  • Sep 18, 2014
  • Abhijit Gosavi
  • Research Article
  • Citations124

Ant-TD: Ant colony optimization plus temporal difference reinforcement learning for multi-label feature selection

  • Apr 27, 2021
  • Swarm and Evolutionary Computation
  • Mohsen Paniri +2
  • Conference Article
  • Citations13

General simulation platform for vision based UAV testing

  • Aug 01, 2015
  • Qing Bu +5
  • Conference Article
  • Citations11

Towards continual reinforcement learning through evolutionary meta-learning

  • Jul 13, 2019
  • Djordje Grbic +1
  • Research Article
  • Citations21

Collaborative Reinforcement Learning Based Unmanned Aerial Vehicle (UAV) Trajectory Design for 3D UAV Tracking

  • Dec 01, 2024
  • IEEE Transactions on Mobile Computing
  • Yujiao Zhu +5
  • Research Article
  • Citations7

A Novel Augmentative Backward Reward Function with Deep Reinforcement Learning for Autonomous UAV Navigation

  • Jul 06, 2022
  • Applied Artificial Intelligence
  • Manit Chansuparp +1
  • Research Article
  • Citations5

Offline Reinforcement Learning for Adaptive Control in Manufacturing Processes: A Press Hardening Case Study

  • Nov 12, 2024
  • Journal of Computing and Information Science in Engineering
  • Nuria Nievas +5
  • PDF
  • Research Article
  • Citations13

Task Offloading Strategy for Unmanned Aerial Vehicle Power Inspection Based on Deep Reinforcement Learning

  • Mar 24, 2024
  • Sensors (Basel, Switzerland)
  • Wei Zhuang +2
  • Research Article
  • Citations15

AQROM: A quality of service aware routing optimization mechanism based on asynchronous advantage actor-critic in software-defined networks

  • Dec 02, 2022
  • Digital Communications and Networks
  • Wei Zhou +6
  • Book Chapter
  • Citations5

Grey Reinforcement Learning for Incomplete Information Processing

  • Jan 01, 2006
  • Chunlin Chen +2
  • PDF
  • Research Article

Bioinspired morphology and task curricula for learning locomotion in bipedal muscle-actuated systems

  • Jun 20, 2025
  • Communications Engineering
  • Nadine Badie +5
  • Conference Article
  • Citations1

Comparison of Reinforcement Learning Methods for Production Control in Discrete Manufacturing Systems

  • Jun 12, 2023
  • Lingxiang Yun +3
  • Conference Article

Analysis about Efficiency of Indirect Media Communication on Multi-agent Cooperation Learning

  • Oct 01, 2006
  • Gang Zhao +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.