论文信息 - Continuous-Time Markov Decision Processes with Controlled Observations

Continuous-Time Markov Decision Processes with Controlled Observations

In this paper, we study a continuous-time discounted jump Markov decision process with both controlled actions and observations. The observation is only available for a discrete set of time instances. At each time of observation, one has to select an optimal timing for the next observation and a control trajectory for the time interval between two observation points. We provide a theoretical framework that the decision maker can utilize to find the optimal observation epochs and the optimal actions jointly. Two cases are investigated. One is gated queueing systems in which we explicitly characterize the optimal action and the optimal observation where the optimal observation is shown to be independent of the state. Another is the inventory control problem with Poisson arrival process in which we obtain numerically the optimal action and observation. The results show that it is optimal to observe more frequently at a region of states where the optimal action adapts constantly.

[1] Abraham Wald,et al. Some Generalizations of the Theory of Cumulative Sums of Random Variables , 1945 .

[2] T. Başar. Minimax control of switching systems under sampling , 1994, Proceedings of 1994 33rd IEEE Conference on Decision and Control.

[3] Jr. Shaler Stidham. Optimal control of admission to a queueing system , 1985 .

[4] Jürgen Pannek,et al. Numerical Optimal Control of Nonlinear Systems , 2011 .

[5] Peter E. Caines,et al. Stochastic optimal control under Poisson-distributed observations , 2000, IEEE Trans. Autom. Control..

[6] John N. Tsitsiklis,et al. Neuro-Dynamic Programming , 1996, Encyclopedia of Machine Learning.

[7] R. Durrett. Probability: Theory and Examples , 1993 .

[8] Martin L. Puterman,et al. Markov Decision Processes: Discrete Stochastic Dynamic Programming , 1994 .

[9] Vikram Krishnamurthy,et al. Partially observed Markov decision processes (POMDPs) , 2016 .

[10] Tamer Basar,et al. Optimal control of LTI systems over unreliable communication links , 2006, Autom..

[11] Eitan Altman,et al. Applications of Markov Decision Processes in Communication Networks , 2000 .

[12] Edwin K. P. Chong,et al. UAV Path Planning in a Dynamic Environment via Partially Observable Markov Decision Process , 2013, IEEE Transactions on Aerospace and Electronic Systems.

[13] Daniel Liberzon,et al. Calculus of Variations and Optimal Control Theory: A Concise Introduction , 2012 .