• Home
  • Search
  • Integrated Double Estimator Architecture for Reinforcement Learning.
  • Cite Icon7
  • https://doi.org/10.1109/tcyb.2020.3023033Copy DOI Icon

Integrated Double Estimator Architecture for Reinforcement Learning.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Estimation bias is an important index for evaluating the performance of reinforcement learning (RL) algorithms. The popular RL algorithms, such as Q -learning and deep Q -network (DQN), often suffer overestimation due to the maximum operation in estimating the maximum expected action values of the next states, while double Q -learning (DQ) and double DQN may fall into underestimation by using a double estimator (DE) to avoid overestimation. To keep the balance between overestimation and underestimation, we propose a novel integrated DE (IDE) architecture by combining the maximum operation and DE operation to estimate the maximum expected action value. Based on IDE, two RL algorithms: 1) integrated DQ (IDQ) and 2) its deep network version, that is, integrated double DQN (IDDQN), are proposed. The main idea of the proposed RL algorithms is that the maximum and DE operations are integrated to eliminate the estimation bias, where one estimator is stochastically used to perform action selection based on the maximum operation, and the convex combination of two estimators is used to carry out action evaluation. We theoretically analyze the reason of estimation bias caused by using nonmaximum operation to estimate the maximum expected value and investigate the possible reasons of underestimation existence in DQ. We also prove the unbiasedness of IDE and convergence of IDQ. Experiments on the grid world and Atari 2600 games indicate that IDQ and IDDQN can reduce or even eliminate estimation bias effectively, enable the learning to be more stable and balanced, and improve the performance effectively.

Similar Papers
  • Research Article
  • Citations1

Optimizing TSCH Scheduling for IIoT Networks Using Reinforcement Learning

  • Sep 03, 2025
  • Technologies
  • Sahar Ben Yaala +2
  • Research Article
  • Citations76

The flying sidekick traveling salesman problem with stochastic travel time: A reinforcement learning approach

  • Jun 28, 2022
  • Transportation Research Part E: Logistics and Transportation Review
  • Zeyu Liu +2
  • Conference Article

On the Design of Safe Continual RL Methods for Control of Nonlinear Systems

  • Jun 24, 2025
  • Austin Coursey +2
  • Research Article
  • Citations23

Deep Reinforcement Learning With Modulated Hebbian Plus Q-Network Architecture.

  • May 01, 2022
  • IEEE Transactions on Neural Networks and Learning Systems
  • Pawel Ladosz +7
  • Research Article
  • Citations5

P2P power trading based on reinforcement learning for nanogrid clusters

  • Jul 19, 2024
  • Expert Systems With Applications
  • Hojun Jin +4
  • PDF
  • Research Article
  • Citations14

Secure State Estimation of Cyber-Physical System under Cyber Attacks: Q-Learning vs. SARSA

  • Oct 01, 2022
  • Electronics
  • Zengwang Jin +5
  • Conference Article
  • Citations12

Solve the inverted pendulum problem base on DQN algorithm

  • Jun 01, 2019
  • Xiaoqian Li +2
  • Research Article
  • Citations4

Enhanced Reinforcement Learning Algorithm Based-Transmission Parameter Selection for Optimization of Energy Consumption and Packet Delivery Ratio in LoRa Wireless Networks

  • Dec 20, 2024
  • Journal of Sensor and Actuator Networks
  • Batyrbek Zholamanov +10
  • Conference Article

Effective Linear Policy Gradient Search through Primal-Dual Approximation

  • Jul 01, 2020
  • Yiming Peng +2
  • Research Article

The Role of Reinforcement Learning in Advancing Artificial Intelligence: An Experimental Study with Q-Learning and DQN

  • Aug 24, 2025
  • The Asian Bulletin of Big Data Management
  • Maryam Gul +5
  • Conference Article
  • Citations7

Constrained Expectation-Maximization Methods for Effective Reinforcement Learning

  • Jul 01, 2018
  • Gang Chen +2
  • Video Transcripts

Active Screening for Recurrent Diseases: A Reinforcement Learning Approach

  • Apr 11, 2021
  • Underline Science Inc.
  • Milind Tambe +3
  • Research Article

Multi‐Agent Reinforcement Learning Algorithm Based on Local Observation Imitation Learning

  • Jan 01, 2025
  • IET Control Theory & Applications
  • Hui Zhang +3
  • Research Article
  • Citations5

Deep Deterministic Policy Gradient to Regulate Feedback Control Systems Using Reinforcement Learning

  • Jan 01, 2022
  • Computers, Materials & Continua
  • Samir Salem Al-Bawri +6
  • PDF
  • Research Article
  • Citations10

A Novel Functional Electrical Stimulation-Induced Cycling Controller Using Reinforcement Learning to Optimize Online Muscle Activation Pattern

  • Nov 24, 2022
  • Sensors
  • Tiago Coelho-Magalhães +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.