论文信息 - Context-Aware Policy Reuse

Context-Aware Policy Reuse

Transfer learning can greatly speed up reinforcement learning for a new task by leveraging policies of relevant tasks. Existing works of policy reuse either focus on only selecting a single best source policy for transfer without considering contexts, or cannot guarantee to learn an optimal policy for a target task. To improve transfer efficiency and guarantee optimality, we develop a novel policy reuse method, called Context-Aware Policy reuSe (CAPS), that enables multi-policy transfer. Our method learns when and which source policy is best for reuse, as well as when to terminate its reuse. CAPS provides theoretical guarantees in convergence and optimality for both source policy selection and target task learning. Empirical results on a grid-based navigation domain and the Pygame Learning Environment demonstrate that CAPS significantly outperforms other state-of-the-art policy reuse methods.

[1] Michael I. Jordan,et al. MASSACHUSETTS INSTITUTE OF TECHNOLOGY ARTIFICIAL INTELLIGENCE LABORATORY and CENTER FOR BIOLOGICAL AND COMPUTATIONAL LEARNING DEPARTMENT OF BRAIN AND COGNITIVE SCIENCES , 1996 .

[2] Yang Gao,et al. Measuring the Distance Between Finite Markov Decision Processes , 2016, AAMAS.

[3] Emilio Soria Olivas,et al. Handbook of Research on Machine Learning Applications and Trends : Algorithms , Methods , and Techniques , 2009 .

[4] Ruslan Salakhutdinov,et al. Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning , 2015, ICLR.

[5] Peter Dayan,et al. Technical Note: Q-Learning , 2004, Machine Learning.

[6] Doina Precup,et al. The Option-Critic Architecture , 2016, AAAI.

[7] S. Shankar Sastry,et al. A Multi-Armed Bandit Approach for Online Expert Selection in Markov Decision Processes , 2017, ArXiv.

[8] Doina Precup,et al. Optimal policy switching algorithms for reinforcement learning , 2010, AAMAS.

[9] Razvan Pascanu,et al. Advances in optimizing recurrent networks , 2012, 2013 IEEE International Conference on Acoustics, Speech and Signal Processing.

[10] Peter Stone,et al. The utility of temporal abstraction in reinforcement learning , 2008, AAMAS.

[11] Eric Eaton,et al. Unsupervised Cross-Domain Transfer in Policy Gradient Reinforcement Learning via Manifold Alignment , 2015, AAAI.

[12] Sergey Levine,et al. Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning , 2017, ICLR.

[13] Tom Schaul,et al. Universal Value Function Approximators , 2015, ICML.

[14] Shie Mannor,et al. Time-Regularized Interrupting Options (TRIO) , 2014, ICML.

[15] Doina Precup,et al. Learning with Options that Terminate Off-Policy , 2017, AAAI.

[16] Richard S. Sutton,et al. Reinforcement Learning: An Introduction , 1998, IEEE Trans. Neural Networks.

[17] Pieter Abbeel,et al. Meta Learning Shared Hierarchies , 2017, ICLR.

[18] Gavriel Salomon,et al. T RANSFER OF LEARNING , 1992 .

[19] Benjamin Rosman,et al. Bayesian policy reuse , 2015, Machine Learning.

[20] Rich Caruana,et al. Multitask Learning , 1997, Machine-mediated learning.

[21] Tie-Yan Liu,et al. Target Transfer Q-Learning and Its Convergence Analysis , 2018, Neurocomputing.