On 2-Armed Gaussian Bandits and Optimization
暂无分享,去创建一个
We explore the 2-armed bandit with Gaussian payoffs as a theoretical model for optimization. We formulate the problem from a Bayesian perspective, and provide the optimal strategy for both 1 and 2 pulls. We present regions of parameter space where a greedy strategy is provably optimal. We also compare the greedy and optimal strategies to a genetic-algorithm-based strategy. In doing so we correct a previous error in the literature concerning the Gaussian bandit problem and the supposed optimality of genetic algorithms for this problem. Finally, we provide an analytically simple bandit model that is more directly applicable to optimization theory than the traditional bandit problem, and determine a near-optimal strategy for that model.
[1] K. Pearson,et al. Biometrika , 1902, The American Naturalist.
[2] John H. Holland,et al. Adaptation in Natural and Artificial Systems: An Introductory Analysis with Applications to Biology, Control, and Artificial Intelligence , 1992 .
[3] P. W. Jones,et al. Bandit Problems, Sequential Allocation of Experiments , 1987 .
[4] D. Wolpert,et al. No Free Lunch Theorems for Search , 1995 .