• Home
  • Search
  • Learning and incentives in user-generated content
  • Cite Icon59
  • https://doi.org/10.1145/2422436.2422465Copy DOI Icon

Learning and incentives in user-generated content

  • Jan 9, 2013
  • Arpita Ghosh +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Motivated by the problem of learning the qualities of user-generated content on the Web, we study a multi-armed bandit problem where the number and success probabilities of the arms of the bandit are endogenously determined by strategic agents in response to the incentives provided by the learning algorithm. We model the contributors of user-generated content as attention-motivated agents who derive benefit when their contribution is displayed, and have a cost to quality, where a contribution's quality is the probability of its receiving a positive viewer vote. Agents strategically choose whether and what quality contribution to produce in response to the algorithm that decides how to display contributions. The algorithm, which would like to eventually only display the highest quality contributions, can only learn a contribution's quality from the viewer votes the contribution receives when displayed. The problem of inferring the relative qualities of contributions using viewer feedback, to optimize for overall viewer satisfaction over time, can then be modeled as the classic multi-armed bandit problem, except that the arms available to the bandit and therefore the achievable regret are endogenously determined by strategic agents --- a good algorithm for this setting must not only quickly identify the best contributions, but also incentivize high-quality contributions to choose amongst in the first place. We first analyze the well-known UCB algorithm Ma [Auer et al. 2002] as a mechanism in this setting, where the total number of potential contributors or arms, K, can grow with the total number of viewers or available periods, T, and the maximum possible success probability of an arm, γ, may be bounded away from 1 to model malicious or error-prone viewers in the audience. We first show that while Ma can incentivize high-quality arms and achieve strong sublinear equilibrium regret when K(T) does not grow too quickly with T, it incentivizes very low quality contributions when K(T) scales proportionally with T. We then show that modifying the UCB mechanism to explore a randomly chosen restricted subset of √{T} arms provides excellent incentive properties --- this modified mechanism achieves strong sublinear regret, which is the regret measured against the maximum achievable quality γ, in every equilibrium, for all ranges of K(T) ≤ T, for all possible values of the audience parameter $\gamma$.

Similar Papers
  • Conference Article
  • Citations82

Thompson Sampling for Dynamic Multi-armed Bandits

  • Dec 01, 2011
  • Neha Gupta +2
  • Research Article
  • Citations1

Robust Performance Incentivizing Algorithms for Multi-Armed Bandits with Strategic Agents

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Seyed A Esmaeili +2
  • Research Article
  • Citations5

Content Contributor Management and Network Effects in a UGC Environment

  • Jan 25, 2011
  • SSRN Electronic Journal
  • Kaifu Zhang +2
  • Research Article

Performance Comparison of UCB, TS, and -Greedy TS Algorithms through Simulation of Multi-Armed Bandit Machine

  • Oct 31, 2024
  • Applied and Computational Engineering
  • Zhuoran Liu
  • Research Article
  • Citations32

Fast Learning for Dynamic Resource Allocation in AI-Enabled Radio Networks

  • Nov 22, 2019
  • IEEE Transactions on Cognitive Communications and Networking
  • Muhammad Anjum Qureshi +1
  • Book Chapter
  • Citations32

Solving Non-Stationary Bandit Problems by Random Sampling from Sibling Kalman Filters

  • Jan 01, 2010
  • Ole-Christoffer Granmo +1
  • Book Chapter
  • Citations5

Successive Reduction of Arms in Multi-Armed Bandits

  • Jan 01, 2011
  • Neha Gupta +2
  • Conference Article
  • Citations6

Exploration. Exploitation, and Engagement in Multi-Armed Bandits with Abandonment

  • Sep 27, 2022
  • Zixian Yang +2
  • PDF
  • Research Article
  • Citations15

Decentralized Heterogeneous Multi-Player Multi-Armed Bandits With Non-Zero Rewards on Collisions

  • Apr 01, 2022
  • IEEE Transactions on Information Theory
  • Akshayaa Magesh +1
  • Supplementary Content

The Properties And Origins Of Spiral Structure Across The Galaxy Population

  • Jul 19, 2018
  • Zenodo (CERN European Organization for Nuclear Research)
  • Ray L Hart
  • Conference Article
  • Citations25

Learning the demand curve in posted-price digital goods auctions

  • May 02, 2011
  • Meenal Chhabra +1
  • Book Chapter
  • Citations85

Multi-Armed Bandit Learning in IoT Networks: Learning Helps Even in Non-stationary Settings

  • Jan 01, 2018
  • Lecture notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering
  • Rémi Bonnefoi +4
  • Research Article
  • Citations8

Quantum Multi-Armed Bandits and Stochastic Linear Bandits Enjoy Logarithmic Regrets

  • Jun 26, 2023
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Zongqi Wan +4
  • Book Chapter
  • Citations28

Bandit problems

  • Apr 25, 2008
  • The New Palgrave Dictionary of Economics
  • Dirk Bergemann +1
  • Conference Article
  • Citations3

Multi-Player Multi-Armed Bandits with Finite Shareable Resources Arms: Learning Algorithms & Applications

  • Jul 01, 2022
  • Xuchuang Wang +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.