A

Optimal data driven resource allocation under multi-armed bandit observations

Annals of Operations Research

Abstract

Abstract This paper introduces the first asymptotically optimal strategy for a multi armed bandit (MAB) model under side constraints. The side constraints model situations in which bandit activations are limited by the availability of certain resources that are replenished at a constant rate. The main result involves the derivation of an asymptotic lower bound for the regret of feasible uniformly fast policies and the construction of policies that achieve this lower bound, under pertinent conditions. Further, we provide the explicit form of such policies for the case in which the unknown distributions are Normal with unknown means and known variances, for the case of Normal distributions with unknown means and unknown variances and for the case of arbitrary discrete distributions with finite support.

Authors 3

  1. National and Kapodistrian University of Athens

    Affiliation as printed

    Department of Mathematics, National and Kapodistrian University of Athens, Athens, Greece

  2. Leiden University

    Affiliation as printed

    Leiden University, Mathematical Institute, Einsteinweg 55, 2333 CC, Leiden, The Netherlands

  3. Rutgers, The State University of New Jersey

    Affiliation as printed

    Department of Management Science and Information Systems, Rutgers University, Piscataway, NJ, 08854, USA

Cited by 3 stored of 3

3 results

No patents citing this paper on Lens.org (checked 2026-10-11).

References 46