• Home
  • Search
  • The Influence of Shape Constraints on the Thresholding Bandit Problem
  • Cite Icon1
  • https://doi.org/10.48550/arxiv.2006.10006Copy DOI Icon

The Influence of Shape Constraints on the Thresholding Bandit Problem

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We investigate the stochastic Thresholding Bandit problem (TBP) under several shape constraints. On top of (i) the vanilla, unstructured TBP, we consider the case where (ii) the sequence of arm's means $(\mu_k)_k$ is monotonically increasing MTBP, (iii) the case where $(\mu_k)_k$ is unimodal UTBP and (iv) the case where $(\mu_k)_k$ is concave CTBP. In the TBP problem the aim is to output, at the end of the sequential game, the set of arms whose means are above a given threshold. The regret is the highest gap between a misclassified arm and the threshold. In the fixed budget setting, we provide problem independent minimax rates for the expected regret in all settings, as well as associated algorithms. We prove that the minimax rates for the regret are (i) $\sqrt{\log(K)K/T}$ for TBP, (ii) $\sqrt{\log(K)/T}$ for MTBP, (iii) $\sqrt{K/T}$ for UTBP and (iv) $\sqrt{\log\log K/T}$ for CTBP, where $K$ is the number of arms and $T$ is the budget. These rates demonstrate that the dependence on $K$ of the minimax regret varies significantly depending on the shape constraint. This highlights the fact that the shape constraints modify fundamentally the nature of the TBP.

Similar Papers
  • Research Article
  • Citations8

The Price of Incentivizing Exploration: A Characterization via Thompson Sampling and Sample Complexity

  • Nov 29, 2022
  • Operations Research
  • Mark Sellke +1
  • Research Article

Performance Comparison of UCB, TS, and -Greedy TS Algorithms through Simulation of Multi-Armed Bandit Machine

  • Oct 31, 2024
  • Applied and Computational Engineering
  • Zhuoran Liu
  • Book Chapter
  • Citations5

Successive Reduction of Arms in Multi-Armed Bandits

  • Jan 01, 2011
  • Neha Gupta +2
  • PDF
  • Research Article
  • Citations15

Decentralized Heterogeneous Multi-Player Multi-Armed Bandits With Non-Zero Rewards on Collisions

  • Apr 01, 2022
  • IEEE Transactions on Information Theory
  • Akshayaa Magesh +1
  • Research Article
  • Citations32

Fast Learning for Dynamic Resource Allocation in AI-Enabled Radio Networks

  • Nov 22, 2019
  • IEEE Transactions on Cognitive Communications and Networking
  • Muhammad Anjum Qureshi +1
  • Conference Article
  • Citations82

Thompson Sampling for Dynamic Multi-armed Bandits

  • Dec 01, 2011
  • Neha Gupta +2
  • Conference Article
  • Citations59

Learning and incentives in user-generated content

  • Jan 09, 2013
  • Arpita Ghosh +1
  • Conference Article
  • Citations11

Approximation Algorithms for Restless Bandit Problems

  • Jan 04, 2009
  • Sudipto Guha +2
  • Research Article

Optimal Best-Arm Identification under Fixed Confidence with Multiple Optima

  • Jan 01, 2026
  • IEEE Transactions on Information Theory
  • Lan Truong
  • Research Article
  • Citations17

Dynamic bargaining game DEA carbon emissions abatement allocation and the Nash equilibrium

  • May 16, 2024
  • Energy Economics
  • Junfei Chu +3
  • Research Article
  • Citations8

Stochastic adaptive dynamical games

  • Jan 01, 2016
  • SCIENTIA SINICA Mathematica
  • Yuan Shuo +1
  • PDF
  • Research Article
  • Citations72

Risk-aware multi-armed bandit problem with application to portfolio selection

  • Nov 01, 2017
  • Royal Society Open Science
  • Xiaoguang Huo +1
  • Book Chapter
  • Citations40

Deviations of Stochastic Bandit Regret

  • Jan 01, 2011
  • Antoine Salomon +1
  • Conference Article
  • Citations68

Bandits with switching costs

  • May 31, 2014
  • Ofer Dekel +3
  • Conference Article
  • Citations10

Mean field equilibria of multi armed bandit games

  • Oct 01, 2012
  • Ramki Gummadi +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.