• Home
  • Search
  • A General Framework for Bandit Problems Beyond Cumulative Objectives
  • Open Access IconOpen Access
  • Cite Icon15
  • https://doi.org/10.1287/moor.2022.1335Copy DOI Icon

A General Framework for Bandit Problems Beyond Cumulative Objectives

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The stochastic multiarmed bandit (MAB) problem is a common model for sequential decision problems. In the standard setup, a decision maker has to choose at every instant between several competing arms; each of them provides a scalar random variable, referred to as a “reward.” Nearly all research on this topic considers the total cumulative reward as the criterion of interest. This work focuses on other natural objectives that cannot be cast as a sum over rewards but rather, more involved functions of the reward stream. Unlike the case of cumulative criteria, in the problems we study here, the oracle policy, which knows the problem parameters a priori and is used to “center” the regret, is not trivial. We provide a systematic approach to such problems and derive general conditions under which the oracle policy is sufficiently tractable to facilitate the design of optimism-based (upper confidence bound) learning policies. These conditions elucidate an interesting interplay between the arm reward distributions and the performance metric. Our main findings are illustrated for several commonly used objectives, such as conditional value-at-risk, mean-variance trade-offs, Sharpe ratio, and more. Funding: This work was partially funded by the Israel Science Foundation [Contract 2199/20] and by the European Community’s Seventh Framework Programme FP7/2007–2013 [Grant 306638 (Scaling Up Reinforcement Learning: Structure Learning, Skill Acquisition, and Reward Shaping)].

Similar Papers
  • Conference Article
  • Citations11

Approximation Algorithms for Restless Bandit Problems

  • Jan 04, 2009
  • Sudipto Guha +2
  • Research Article

Performance Comparison of UCB, TS, and -Greedy TS Algorithms through Simulation of Multi-Armed Bandit Machine

  • Oct 31, 2024
  • Applied and Computational Engineering
  • Zhuoran Liu
  • Conference Article
  • Citations1

Application of the UCT Algorithm for Noisy Optimization Problems

  • Aug 01, 2016
  • Akira Notsu +3
  • Research Article
  • Citations2

From Learning to Analytics: Improving Model Efficacy With Goal-Directed Client Selection

  • Dec 01, 2024
  • IEEE Transactions on Mobile Computing
  • Jingwen Tong +4
  • Research Article
  • Citations25

Gap-free Bounds for Stochastic Multi-Armed Bandit

  • Jan 01, 2008
  • IFAC Proceedings Volumes
  • A Juditsky +3
  • Conference Article

Optimization of Transmission Strategy for Wireless Power Transfer Using Multi-Armed Bandit Algorithm

  • Oct 27, 2021
  • Yuan Xing +5
  • Research Article

Multi-Armed Bandits: Algorithms, Applications, and Future Directions

  • Jan 29, 2026
  • Academic Journal of Science and Technology
  • Yurong Zheng
  • Research Article
  • Citations1

Block pruning residual networks using Multi-Armed Bandits

  • Dec 07, 2024
  • Journal of Experimental & Theoretical Artificial Intelligence
  • Mohamed Akrem Benatia +3
  • Research Article

Modeling user autonomy in recommender systems using Markov perturbation-based multi-armed bandits

  • Feb 21, 2025
  • Theoretical and Natural Science
  • Yuxuan Yang
  • Research Article

Scaling and Optimizing Consumer Tech Products with Multi-Armed Bandit Algorithms: Applications in eCommerce

  • Mar 03, 2025
  • International Journal of Scientific Research in Computer Science, Engineering and Information Technology
  • Siddharth Gupta
  • Research Article

Utilizing Multi-Armed Bandit Algorithms for Advertising: An In-Depth Case Study on an Online Retail Platform's Advertising Campaign

  • Apr 26, 2024
  • Highlights in Science, Engineering and Technology
  • Litian Shao
  • Research Article
  • Citations3

Generalized Bandits with Learning and Queueing in Split Liver Transplantation

  • Jan 01, 2021
  • SSRN Electronic Journal
  • Yanhan Tang +2
  • Book Chapter
  • Citations40

Deviations of Stochastic Bandit Regret

  • Jan 01, 2011
  • Antoine Salomon +1
  • Research Article

User Opinion based Trust Value Prediction for Online Social Network

  • Aug 15, 2020
  • International Journal of Scientific Research in Computer Science, Engineering and Information Technology
  • Abinaya R
  • Conference Article
  • Citations7

Aging Wireless Bandits: Regret Analysis and Order-Optimal Learning Algorithm

  • Oct 18, 2021
  • Eray Unsal Atay +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.