论文信息 - Markov decision processes under ambiguity

Markov decision processes under ambiguity

We consider statistical Markov Decision Processes where the decision maker is risk averse against model ambiguity. The latter is given by an unknown parameter which influences the transition law and the cost functions. Risk aversion is either measured by the entropic risk measure or by the Average Value at Risk. We show how to solve these kind of problems using a general minimax theorem. Under some continuity and compactness assumptions we prove the existence of an optimal (deterministic) policy and discuss its computation. We illustrate our results using an example from statistical decision theory.

U. Rieder | Nicole Bauerle

[1] K. Hinderer,et al. Foundations of Non-stationary Dynamic Programming with Discrete Time Parameter , 1970 .

[2] M. Schäl. On dynamic programming: Compactness of the space of policies , 1975 .

[3] H. Föllmer,et al. Stochastic Finance: An Introduction in Discrete Time , 2002 .

[4] A. Bhattacharyya. On a measure of divergence between two statistical populations defined by their probability distributions , 1943 .

[5] Nicole Bäuerle,et al. Partially Observable Risk-Sensitive Markov Decision Processes , 2015, Math. Oper. Res..

[6] M. Sion. On general minimax theorems , 1958 .

[7] M. Degroot. Optimal Statistical Decisions , 1970 .

[8] M. Schal. On Dynamic Programming and Statistical Decision Theory , 1979 .

[9] Takayuki Osogami,et al. Robustness and risk-sensitivity in Markov decision processes , 2012, NIPS.

[10] A. Rustichini,et al. Ambiguity Aversion, Robustness, and the Variational Representation of Preferences , 2006 .

[11] Garud Iyengar,et al. Robust Dynamic Programming , 2005, Math. Oper. Res..