Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > sci.physics > #521756
| Subject | Re: Single photon decision-maker solves multi-armed bandit problem |
|---|---|
| Newsgroups | sci.physics |
| References | <_7WdnSXU8PhSsmbInZ2dnUU7-QNi4p2d@giganews.com> <9uorcc-ph1.ln1@mail.specsol.com> |
| From | Sam Wormley <swormley1@gmail.com> |
| Date | 2015-09-17 21:07 -0500 |
| Message-ID | <_7WdnSPU8PhS7GbInZ2dnUU7-QMAAAAA@giganews.com> (permalink) |
Review for the jimp | Multi-armed bandit > https://en.wikipedia.org/wiki/Multi-armed_bandit > In probability theory, the multi-armed bandit problem (sometimes > called the K-[1] or N-armed bandit problem[2]) is a problem in which > a gambler at a row of slot machines (sometimes known as "one-armed > bandits") has to decide which machines to play, how many times to > play each machine and in which order to play them.[3] When played, > each machine provides a random reward from a distribution specific to > that machine. The objective of the gambler is to maximize the sum of > rewards earned through a sequence of lever pulls.[4][5] > > Robbins in 1952, realizing the importance of the problem, constructed > convergent population selection strategies in "some aspects of the > sequential design of experiments".[6] > > A theorem, the Gittins index published first by John C. Gittins gives > an optimal policy in the Markov setting for maximizing the expected > discounted reward.[7] > > In practice, multi-armed bandits have been used to model the problem > of managing research projects in a large organization, like a science > foundation or a pharmaceutical company. Given a fixed budget, the > problem is to allocate resources among the competing projects, whose > properties are only partially known at the time of allocation, but > which may become better understood as time passes.[4][5] > > In early versions of the multi-armed bandit problem, the gambler has > no initial knowledge about the machines. The crucial tradeoff the > gambler faces at each trial is between "exploitation" of the machine > that has the highest expected payoff and "exploration" to get more > information about the expected payoffs of the other machines. The > trade-off between exploration and exploitation is also faced in > reinforcement learning. -- sci.physics is an unmoderated newsgroup dedicated to the discussion of physics, news from the physics community, and physics-related social issues.
Back to sci.physics | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Single photon decision-maker solves multi-armed bandit problem Sam Wormley <swormley1@gmail.com> - 2015-09-17 16:25 -0500
Re: Single photon decision-maker solves multi-armed bandit problem jimp@specsol.spam.sux.com - 2015-09-17 21:46 +0000
Re: Single photon decision-maker solves multi-armed bandit problem Sam Wormley <swormley1@gmail.com> - 2015-09-17 21:07 -0500
Re: Single photon decision-maker solves multi-armed bandit problem jimp@specsol.spam.sux.com - 2015-09-18 03:14 +0000
Re: The ass hat spams some more jimp@specsol.spam.sux.com - 2015-09-18 17:46 +0000
Re: Single photon decision-maker solves multi-armed bandit problem "reber g=emc^2" <herbertglazier0@gmail.com> - 2015-09-18 11:43 -0700
Re: Single photon decision-maker solves multi-armed bandit problem benj <nobody@gmail.com> - 2015-09-18 14:48 -0400
Re: Single photon decision-maker solves multi-armed bandit problem Double-A <double-a3@hush.com> - 2015-09-18 16:29 -0700
Re: Single photon decision-maker solves multi-armed bandit problem "reber g=emc^2" <herbertglazier0@gmail.com> - 2015-09-20 09:47 -0700
Re: Single photon decision-maker solves multi-armed bandit problem "Y.Porat" <y.y.porat@gmail.com> - 2015-09-19 01:55 -0700
Re: Single photon decision-maker solves multi-armed bandit problem "Oliver Jacob Bacon" <invalid@example.com> - 2015-09-19 07:58 -0700
Re: Single photon decision-maker solves multi-armed bandit problem "Y.Porat" <y.y.porat@gmail.com> - 2015-09-19 01:57 -0700
Re: Single photon decision-maker solves multi-armed bandit problem "Oscar Jazz Bottom" <invalid@example.com> - 2015-09-19 07:57 -0700
Re: Single photon decision-maker solves multi-armed bandit problem "Y.Porat" <y.y.porat@gmail.com> - 2015-09-20 01:21 -0700
csiph-web