论文信息 - Self-Imitation Learning via Trajectory-Conditioned Policy for Hard-Exploration Tasks

Self-Imitation Learning via Trajectory-Conditioned Policy for Hard-Exploration Tasks

Imitation learning from human-expert demonstrations has been shown to be greatly helpful for challenging reinforcement learning problems with sparse environment rewards. However, it is very difficult to achieve similar success without relying on expert demonstrations. Recent works on self-imitation learning showed that imitating the agent's own past good experience could indirectly drive exploration in some environments, but these methods often lead to sub-optimal and myopic behavior. To address this issue, we argue that exploration in diverse directions by imitating diverse trajectories, instead of focusing on limited good trajectories, is more desirable for the hard-exploration tasks. We propose a new method of learning a trajectory-conditioned policy to imitate diverse trajectories from the agent's own past experience and show that such self-imitation helps avoid myopic behavior and increases the chance of finding a globally optimal solution for hard-exploration tasks, especially when there are misleading rewards. Our method significantly outperforms existing self-imitation learning and count-based exploration methods on various hard-exploration tasks with local optima. In particular, we report a state-of-the-art score of more than 20,000 points on Montezuma's Revenge without using expert demonstrations or resetting to arbitrary states.

[1] Honglak Lee,et al. Contingency-Aware Exploration in Reinforcement Learning , 2018, ICLR.

[2] Kenneth O. Stanley,et al. Go-Explore: a New Approach for Hard-Exploration Problems , 2019, ArXiv.

[3] Peter Auer,et al. Using Confidence Bounds for Exploitation-Exploration Trade-offs , 2003, J. Mach. Learn. Res..

[4] Sebastian Thrun,et al. Active Exploration in Dynamic Environments , 1991, NIPS.

[5] Tom Schaul,et al. Prioritized Experience Replay , 2015, ICLR.

[6] Satinder Singh,et al. Self-Imitation Learning , 2018, ICML.

[7] A. P. Hyper-parameters. Count-Based Exploration with Neural Density Models , 2017 .

[8] Yoshua Bengio,et al. Neural Machine Translation by Jointly Learning to Align and Translate , 2014, ICLR.

[9] Shane Legg,et al. Human-level control through deep reinforcement learning , 2015, Nature.

[10] Marlos C. Machado,et al. Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents , 2017, J. Artif. Intell. Res..

[11] Stefanie Tellex,et al. Deep Abstract Q-Networks , 2017, AAMAS.

[12] Jürgen Schmidhuber,et al. Curious model-building control systems , 1991, [Proceedings] 1991 IEEE International Joint Conference on Neural Networks.

[13] Jitendra Malik,et al. Combining self-supervised learning and imitation for vision-based rope manipulation , 2017, 2017 IEEE International Conference on Robotics and Automation (ICRA).

[14] Michael L. Littman,et al. An analysis of model-based Interval Estimation for Markov Decision Processes , 2008, J. Comput. Syst. Sci..

[15] Sergey Levine,et al. Divide-and-Conquer Reinforcement Learning , 2017, ICLR.

[16] Richard Tanburn,et al. Making Efficient Use of Demonstrations to Solve Hard Exploration Problems , 2019, ICLR.

[17] Marcin Andrychowicz,et al. Hindsight Experience Replay , 2017, NIPS.

[18] Sergey Levine,et al. Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models , 2015, ArXiv.

[19] Pieter Abbeel,et al. Benchmarking Deep Reinforcement Learning for Continuous Control , 2016, ICML.

[20] J. Urgen Schmidhuber,et al. Adaptive confidence and adaptive curiosity , 1991, Forschungsberichte, TU Munich.

[21] Yoshua Bengio,et al. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation , 2014, EMNLP.

[22] Satinder Singh,et al. Generative Adversarial Self-Imitation Learning , 2018, ArXiv.

[23] Amos J. Storkey,et al. Exploration by Random Network Distillation , 2018, ICLR.

[24] Yishay Mansour,et al. Policy Gradient Methods for Reinforcement Learning with Function Approximation , 1999, NIPS.