• Home
  • Search
  • A novel multi-agent dynamic portfolio optimization learning system based on hierarchical deep reinforcement learning
  • Cite Icon11
  • https://doi.org/10.1007/s40747-025-01884-yCopy DOI Icon

A novel multi-agent dynamic portfolio optimization learning system based on hierarchical deep reinforcement learning

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Deep reinforcement learning (DRL) has been extensively used to address portfolio optimization problems. DRL agents acquire knowledge and make decisions through unsupervised interactions with their environment without requiring explicit knowledge of the joint dynamics of portfolio assets. Among these DRL algorithms, the combination of actor-critic algorithms and deep function approximators is the most widely used DRL algorithm. Here, we find that training the DRL agent using the actor-critic algorithm and deep function approximators may lead to scenarios where the improvement in the DRL agent's risk-adjusted profitability is insignificant. We argue that such situations primarily arise from the following two problems: sparsity in positive reward and the curse of dimensionality. These limitations prevent DRL agents from comprehensively learning asset price change patterns in the training environment. As a result, the DRL agents cannot effectively explore the dynamic portfolio optimization policy to improve the risk-adjusted profitability in the training process. To address these problems, we propose a novel multi-agent learning system based on the hierarchical deep reinforcement learning (HDRL) algorithmic framework in this research. Under this framework, the agents work together as a learning system for portfolio optimization. Specifically, by designing an auxiliary agent that works together with the executive agent for optimal policy exploration, the learning system can focus on exploring the policy with higher risk-adjusted return in the action space with positive return and low variance. The performance of the proposed learning system is evaluated using a portfolio of 29 stocks from the Dow Jones index in four different experiments. In the training process, the objective functions of the actor and critic both ultimately achieve stable convergence in the training process. The risk-adjusted profitability of our learning system in the training environment is significantly improved. Hence, we prove that the policies executed by our learning system in out-sample experiments originate from the DRL agents' comprehensive learning of asset price change patterns in the training environment. Furthermore, we find that adopting the auxiliary agent and HDRL training algorithm can efficiently overcome the issue of the curse of dimensionality and improve the training efficiency in the positive reward sparse environment. In each back-test experiment, the proposed learning system is compared to sixteen traditional strategies and ten strategies based on machine learning algorithms in the performance of profitability and risk control ability. The empirical results in the four evaluation experiments demonstrate the efficacy of our learning system, which outperforms all other strategies by at least 8.2% in terms of Sharpe ratio, Sorino ratio, and Calmar ratio. This indicates that the policies learned in the training environment can exhibit excellent generalization ability in the back-testing experiments.

Similar Papers
  • Research Article
  • Citations25

Security-Aware Resource Allocation Scheme Based on DRL in Cloud–Edge–Terminal Cooperative Vehicular Network

  • Jan 01, 2024
  • IEEE Internet of Things Journal
  • Yi Zhang +2
  • PDF
  • Research Article
  • Citations139

Energy Management of Smart Home with Home Appliances, Energy Storage System and Electric Vehicle: A Hierarchical Deep Reinforcement Learning Approach

  • Apr 10, 2020
  • Sensors (Basel, Switzerland)
  • Sangyoon Lee +1
  • PDF
  • Research Article
  • Citations67

A Hierarchical Deep Reinforcement Learning Framework With High Efficiency and Generalization for Fast and Safe Navigation

  • May 01, 2023
  • IEEE Transactions on Industrial Electronics
  • Wei Zhu +1
  • Conference Article
  • Citations36

An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents

  • Aug 01, 2019
  • Felipe Petroski Such +10
  • Book Chapter
  • Citations9

A Study on Dense and Sparse (Visual) Rewards in Robot Policy Learning

  • Jan 01, 2021
  • Abdalkarim Mohtasib +2
  • PDF
  • Research Article
  • Citations33

Reinforcement Learning-Based School Energy Management System

  • Dec 01, 2020
  • Energies
  • Yassine Chemingui +2
  • Research Article
  • Citations2

Multi‐Objective Bayesian Optimization of Deep Reinforcement Learning for Environmental, Social, and Governance (ESG) Financial Portfolio Management

  • Jun 01, 2025
  • Intelligent Systems in Accounting, Finance and Management
  • Eduardo C Garrido‐Merchán +2
  • Research Article
  • Citations38

Deep reinforcement learning for conservation decisions

  • Sep 14, 2022
  • Methods in Ecology and Evolution
  • Marcus Lapeyrolerie +3
  • Conference Article
  • Citations16

Learning to Play General Video-Games via an Object Embedding Network

  • Aug 01, 2018
  • William Woof +1
  • Conference Article
  • Citations5

DREVAN: Deep Reinforcement Learning-based Vulnerability-Aware Network Adaptations for Resilient Networks

  • Oct 04, 2021
  • Qisheng Zhang +3
  • Research Article
  • Citations19

Multi-Agent and Cooperative Deep Reinforcement Learning for Scalable Network Automation in Multi-Domain SD-EONs

  • Dec 01, 2021
  • IEEE Transactions on Network and Service Management
  • Baojia Li +3
  • Research Article
  • Citations13

ReCARL: Resource Allocation in Cloud RANs with Deep Reinforcement Learning

  • Jan 01, 2020
  • IEEE Transactions on Mobile Computing
  • Zhiyuan Xu +6
  • PDF
  • Research Article
  • Citations37

Deep Reinforcement Learning-Based Smart Joint Control Scheme for On/Off Pumping Systems in Wastewater Treatment Plants

  • Jan 01, 2021
  • IEEE Access
  • Giup Seo +4
  • Research Article
  • Citations2

Deep Reinforcement Learning‐Based Control for Real‐Time Hybrid Simulation of Civil Structures

  • Jan 25, 2025
  • International Journal of Robust and Nonlinear Control
  • Andrés Felipe Niño +6
  • Conference Article
  • Citations8

Energy-efficient train control method based on soft actor-critic algorithm

  • Sep 19, 2021
  • Q Zhu +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.