论文信息 - Perturbations for Adaptive Regulation and Learning

Perturbations for Adaptive Regulation and Learning

Design of adaptive algorithms for simultaneous regulation and estimation of MIMO linear dynamical systems is a canonical reinforcement learning problem. Efficient policies whose regret (i.e. increase in the cost due to uncertainty) scales at a squareroot rate of time have been studied extensively in the recent literature. Nevertheless, existing strategies are computationally intractable and require a priori knowledge of key system parameters. The only exception is a randomized Greedy regulator, for which asymptotic regret bounds have been recently established. However, randomized Greedy leads to probable fluctuations in the trajectory of the system, which renders its finite time regret suboptimal. This work addresses the above issues by designing policies that utilize input signals perturbations. We show that perturbed Greedy guarantees non-asymptotic regret bounds of (nearly) square-root magnitude w.r.t. time. More generally, we establish high probability bounds on both the regret and the learning accuracy under arbitrary input perturbations. The settings where Greedy attains the information theoretic lower bound of logarithmic regret are also discussed. To obtain the results, stateof-the-art tools from martingale theory together with the recently introduced method of policy decomposition are leveraged. Beside adaptive regulators, analysis of input perturbations captures key applications including remote sensing and distributed control.

Mohamad Kazem Shirani Faradonbeh | Ambuj Tewari | G. Michailidis

[1] Ambuj Tewari,et al. On Optimality of Adaptive Linear-Quadratic Regulators , 2018, ArXiv.

[2] T. Lai,et al. Least Squares Estimates in Stochastic Regression Models with Applications to Identification and Control of Dynamic Systems , 1982 .

[3] Sean P. Meyn. Control Techniques for Complex Networks: Workload , 2007 .

[4] P. Kumar,et al. Convergence of adaptive control schemes using least-squares parameter estimates , 1990 .

[5] Joel A. Tropp,et al. User-Friendly Tail Bounds for Sums of Random Matrices , 2010, Found. Comput. Math..

[6] T. Lai,et al. Parallel recursive algorithms in asymptotically efficient adaptive control of linear stochastic systems , 1991 .

[7] Ambuj Tewari,et al. Optimistic Linear Programming gives Logarithmic Regret for Irreducible MDPs , 2007, NIPS.

[8] Ambuj Tewari,et al. Optimality of Fast-Matching Algorithms for Random Networks With Applications to Structural Controllability , 2015, IEEE Transactions on Control of Network Systems.

[9] T. Söderström. Discrete-Time Stochastic Systems: Estimation and Control , 1995 .

[10] Sean P. Meyn,et al. Distributed Control Design for Balancing the Grid Using Flexible Loads , 2018 .

[11] T. Lai,et al. Extended least squares and their applications to adaptive control and prediction in linear systems , 1986 .