• Home
  • Search
  • Diversity-Driven Model Ensemble Adaptive Trust Region Policy Optimization
  • https://doi.org/10.1109/tsmc.2026.3650850Copy DOI Icon

Diversity-Driven Model Ensemble Adaptive Trust Region Policy Optimization

  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Model-based reinforcement learning (MBRL) aims to promote sample efficiency and reduce the number of interactions with the true environment, via learning an environment dynamic model, compared with model-free reinforcement learning (MFRL). However, the success of MBRL heavily relies on two key aspects: model learning and planning. The former refers to learning an accurate model, and the latter aims to improve the behavior policy. In this article, we investigate these two aspects further with model ensemble learning. We design a deep residual attention U-Net (RauNet) with fewer neurons (or weights) than the widely used shallow neural network as our base models and further apply the Hilbert–Schmidt independence criterion (HSIC) as a regularization term to pursue model diversity explicitly for the model ensemble. Furthermore, we propose an adaptive trust region policy optimization (TRPO), in which the parametric Rényi alpha divergence substitutes for the Kullback–Leibler (KL) divergence for measuring the difference between two successive policies, and the alpha value can be adaptively adjusted during TRPO training iterations. This method is called diversity-driven model ensemble adaptive TRPO, or simply diversity-driven model ensemble adaptive trust region policy optimization. Our detailed experiments on six benchmark environments show that our proposed approach is optimal, compared with five state-of-the-art RL techniques.

Similar Papers
  • Research Article
  • Citations39

Navigating complex decision spaces: Problems and paradigms in sequential choice.

  • Jan 01, 2014
  • Psychological Bulletin
  • Matthew M Walsh +1
  • Conference Article
  • Citations2

Context-dependent meta-control for reinforcement learning using a Dirichlet process Gaussian mixture model

  • Jan 01, 2018
  • Dongjae Kim +1
  • PDF
  • Research Article
  • Citations7

Improving Model-Based Deep Reinforcement Learning with Learning Degree Networks and Its Application in Robot Control

  • Mar 04, 2022
  • Journal of Robotics
  • Guoqing Ma +3
  • PDF
  • Research Article
  • Citations32

Parallel model-based and model-free reinforcement learning for card sorting performance

  • Sep 22, 2020
  • Scientific Reports
  • Alexander Steinke +2
  • Peer Review Report

Reviewer #2 (Public review): Neural signatures of model-based and model-free reinforcement learning across prefrontal cortex and striatum

  • Feb 27, 2026
  • Bruno Miranda +5
  • Conference Article
  • Citations1

Hierarchical Control Architecture Regulating Competition between Model-Based and Context-Dependent Model-Free Reinforcement Learning Strategies

  • Oct 01, 2018
  • Dongjae Kim +2
  • Conference Article
  • Citations19

An Overview of Robust Reinforcement Learning

  • Oct 30, 2020
  • Shiyu Chen +1
  • Research Article
  • Citations185

Distributed Coding of Actual and Hypothetical Outcomes in the Orbital and Dorsolateral Prefrontal Cortex

  • May 01, 2011
  • Neuron
  • Hiroshi Abe +1
  • Research Article
  • Citations321

Habits, action sequences and reinforcement learning

  • Apr 01, 2012
  • European Journal of Neuroscience
  • Amir Dezfouli +1
  • Research Article
  • Citations2

Military Decision Support with Actor and Critic Reinforcement Learning Agents

  • Feb 26, 2024
  • Defence Science Journal
  • Jungmok Ma
  • Components

Effects of subclinical depression on prefrontal–striatal model-based and model-free learning

  • May 14, 2021
  • Samuel J Gershman +4
  • Research Article
  • Citations12

Fuzzy-based predictive deep reinforcement learning for robust and constrained optimal control of industrial solar thermal plants

  • Feb 24, 2024
  • Applied Soft Computing
  • Fitsum Bekele Tilahun
  • PDF
  • Supplementary Content
  • Citations3

Understanding cingulotomy’s therapeutic effect in OCD through computer models

  • Jan 10, 2023
  • Frontiers in Integrative Neuroscience
  • Mohamed A Sherif +3
  • Conference Article
  • Citations4

Constrained Policy Optimization Algorithm for Autonomous Driving via Reinforcement Learning

  • Jul 23, 2021
  • Qi Kong +2
  • Book Chapter

Reinforcement Learning Approaches for Optimal Autonomous System Performance

  • Aug 23, 2024
  • B C Shreedevi
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.