• Home
  • Search
  • LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
  • https://doi.org/10.1609/aaai.v40i30.39689Copy DOI Icon

LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration

  • Mar 14, 2026
  • Ruiyu Qiu +4 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Lexicographic multi-objective problems, which consist of multiple conflicting subtasks with explicit priorities, are common in real-world applications. Despite the advantages of Reinforcement Learning (RL) in single tasks, extending conventional RL methods to prioritized multiple objectives remains challenging. In particular, traditional Safe RL and Multi-Objective RL (MORL) methods have difficulty enforcing priority orderings efficiently. Therefore, Lexicographic Multi-Objective RL (LMORL) methods have been developed to address these challenges. However, existing LMORL methods either rely on heuristic threshold tuning with prior knowledge or are restricted to discrete domains. To overcome these limitations, we propose Lexicographically Projected Policy Gradient RL (LPPG-RL), a novel LMORL framework which leverages sequential gradient projections to identify feasible policy update directions, thereby enabling LPPG-RL broadly compatible with all policy gradient algorithms in continuous spaces. LPPG-RL reformulates the projection step as an optimization problem, and utilizes Dykstra's projection rather than generic solvers to deliver great speedups, especially for small- to medium-scale instances. In addition, LPPG-RL introduces Subproblem Exploration (SE) to prevent gradient vanishing, accelerate convergence and enhance stability. We provide theoretical guarantees for convergence and establish a lower bound on policy improvement. Finally, through extensive experiments in a 2D navigation environment, we demonstrate the effectiveness of LPPG-RL, showing that it outperforms existing state-of-the-art continuous LMORL methods.

Similar Papers
  • Research Article
  • Citations13

Reinforcement Learning for Clinical Applications.

  • Feb 08, 2023
  • Clinical Journal of the American Society of Nephrology
  • Kia Khezeli +5
  • Book Chapter
  • Citations11

Approximate Dynamic Programming

  • Feb 28, 2013
  • Rémi Munos
  • Conference Article
  • Citations1

Comparison of Reinforcement Learning Methods for Production Control in Discrete Manufacturing Systems

  • Jun 12, 2023
  • Lingxiang Yun +3
  • Conference Article
  • Citations15

Cola-HRL: Continuous-Lattice Hierarchical Reinforcement Learning for Autonomous Driving

  • Oct 23, 2022
  • Lingping Gao +7
  • Conference Article

QMDP: DASH Adaptation using Queueing Theory within a Markov Decision Process

  • Jan 09, 2021
  • Kevin Gatimu +1
  • Research Article
  • Citations5

5G Network Slicing: Methods to Support Blockchain and Reinforcement Learning

  • Mar 24, 2022
  • Computational Intelligence and Neuroscience
  • Juan Hu +1
  • Research Article
  • Citations31

The reinforcement learning method for occupant behavior in building control: A review

  • Sep 02, 2020
  • Energy and Built Environment
  • Mengjie Han +4
  • Research Article
  • Citations40

A bi-objective deep reinforcement learning approach for low-carbon-emission high-speed railway alignment design

  • Dec 31, 2022
  • Transportation Research Part C: Emerging Technologies
  • Qing He +7
  • Supplementary Content

Sample-Efficient I-Projections for Robot Learning

  • Apr 19, 2021
  • TUbilio (Technical University of Darmstadt)
  • Oleg Arenz
  • PDF
  • Research Article
  • Citations5

A Vision-Based End-to-End Reinforcement Learning Framework for Drone Target Tracking

  • Oct 30, 2024
  • Drones
  • Xun Zhao +4
  • Conference Article
  • Citations5

Cohesion-driven Online Actor-Critic Reinforcement Learning for mHealth Intervention

  • Aug 15, 2018
  • Feiyun Zhu +4
  • Research Article
  • Citations2

Controlling Cable Driven Parallel Robots Operations—Deep Reinforcement Learning Approach

  • Jan 01, 2025
  • IEEE Access
  • Muhammad Kamran Joyo +6
  • Research Article
  • Citations122

AUV position tracking and trajectory control based on fast-deployed deep reinforcement learning method

  • Dec 31, 2021
  • Ocean Engineering
  • Yuan Fang +3
  • PDF
  • Research Article
  • Citations12

Study on Optimization Design of Airfoil Transonic Buffet with Reinforcement Learning Method

  • May 20, 2023
  • Aerospace
  • Hao Chen +4
  • Research Article
  • Citations12

Fuzzy-based predictive deep reinforcement learning for robust and constrained optimal control of industrial solar thermal plants

  • Feb 24, 2024
  • Applied Soft Computing
  • Fitsum Bekele Tilahun
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.