论文信息 - On the use of Deep Autoencoders for Efficient Embedded Reinforcement Learning

On the use of Deep Autoencoders for Efficient Embedded Reinforcement Learning

In autonomous embedded systems, it is often vital to reduce the amount of actions taken in the real world and energy required to learn a policy. Training reinforcement learning agents from high dimensional image representations can be very expensive and time consuming. Autoencoders are deep neural network used to compress high dimensional data such as pixelated images into small latent representations. This compression model is vital to efficiently learn policies, especially when learning on embedded systems. We have implemented this model on the NVIDIA Jetson TX2 embedded GPU, and evaluated the power consumption, throughput, and energy consumption of the autoencoders for various CPU/GPU core combinations, frequencies, and model parameters. Additionally, we have shown the reconstructions generated by the autoencoder to analyze the quality of the generated compressed representation and also the performance of the reinforcement learning agent. Finally, we have presented an assessment of the viability of training these models on embedded systems and their usefulness in developing autonomous policies. Using autoencoders, we were able to achieve 4-5X improved performance compared to a baseline RL agent with a convolutional feature extractor, while using less than 2W of power.

[1] Alex Graves,et al. Asynchronous Methods for Deep Reinforcement Learning , 2016, ICML.

[2] Alec Radford,et al. Proximal Policy Optimization Algorithms , 2017, ArXiv.

[3] Peter Stone,et al. Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces , 2017, AAAI.

[4] Yishay Mansour,et al. Policy Gradient Methods for Reinforcement Learning with Function Approximation , 1999, NIPS.

[5] Tim Oates,et al. SensorNet: A Scalable and Low-Power Deep Convolutional Neural Network for Multimodal Data Classification , 2019, IEEE Transactions on Circuits and Systems I: Regular Papers.

[6] Shane Legg,et al. Human-level control through deep reinforcement learning , 2015, Nature.

[7] Vinicius G. Goecks,et al. Cycle-of-Learning for Autonomous Systems from Human Interaction , 2018, ArXiv.

[8] Owain Evans,et al. Trial without Error: Towards Safe Reinforcement Learning via Human Intervention , 2017, AAMAS.

[9] Sergey Levine,et al. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor , 2018, ICML.