论文信息 - Two-stage training algorithm for AI robot soccer

Two-stage training algorithm for AI robot soccer

In multi-agent reinforcement learning, the cooperative learning behavior of agents is very important. In the field of heterogeneous multi-agent reinforcement learning, cooperative behavior among different types of agents in a group is pursued. Learning a joint-action set during centralized training is an attractive way to obtain such cooperative behavior; however, this method brings limited learning performance with heterogeneous agents. To improve the learning performance of heterogeneous agents during centralized training, two-stage heterogeneous centralized training which allows the training of multiple roles of heterogeneous agents is proposed. During training, two training processes are conducted in a series. One of the two stages is to attempt training each agent according to its role, aiming at the maximization of individual role rewards. The other is for training the agents as a whole to make them learn cooperative behaviors while attempting to maximize shared collective rewards, e.g., team rewards. Because these two training processes are conducted in a series in every time step, agents can learn how to maximize role rewards and team rewards simultaneously. The proposed method is applied to 5 versus 5 AI robot soccer for validation. The experiments are performed in a robot soccer environment using Webots robot simulation software. Simulation results show that the proposed method can train the robots of the robot soccer team effectively, achieving higher role rewards and higher team rewards as compared to other three approaches that can be used to solve problems of training cooperative multi-agent. Quantitatively, a team trained by the proposed method improves the score concede rate by 5% to 30% when compared to teams trained with the other approaches in matches against evaluation teams.

[1] Zihan Zhou,et al. CityFlow: A Multi-Agent Reinforcement Learning Environment for Large Scale City Traffic Scenario , 2019, WWW.

[2] Demis Hassabis,et al. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play , 2018, Science.

[3] Tom Schaul,et al. Dueling Network Architectures for Deep Reinforcement Learning , 2015, ICML.

[4] Karl Tuyls,et al. Markov Security Games : Learning in Spatial Security Problems , 2016 .

[5] Taeyoung Kim,et al. Batch Prioritization in Multigoal Reinforcement Learning , 2020, IEEE Access.

[6] Guy Lever,et al. Emergent Coordination Through Competition , 2019, ICLR.

[7] Olivier Michel,et al. Cyberbotics Ltd. Webots™: Professional Mobile Robot Simulation , 2004 .

[8] Demis Hassabis,et al. Mastering the game of Go without human knowledge , 2017, Nature.

[9] Luiz Felipe Vecchietti,et al. Sampling Rate Decay in Hindsight Experience Replay for Robot Control , 2020, IEEE Transactions on Cybernetics.

[10] Quoc V. Le,et al. HyperNetworks , 2016, ICLR.

[11] Alex Graves,et al. Playing Atari with Deep Reinforcement Learning , 2013, ArXiv.

[12] Wojciech M. Czarnecki,et al. Grandmaster level in StarCraft II using multi-agent reinforcement learning , 2019, Nature.

[13] Yoshua Bengio,et al. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling , 2014, ArXiv.

[14] Demis Hassabis,et al. Mastering the game of Go with deep neural networks and tree search , 2016, Nature.

[15] Yi Wu,et al. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments , 2017, NIPS.

[16] Dongsoo Har,et al. Rewards Prediction-Based Credit Assignment for Reinforcement Learning With Sparse Binary Rewards , 2019, IEEE Access.

[17] Shimon Whiteson,et al. The StarCraft Multi-Agent Challenge , 2019, AAMAS.

[18] Marcin Andrychowicz,et al. Sim-to-Real Transfer of Robotic Control with Dynamics Randomization , 2017, 2018 IEEE International Conference on Robotics and Automation (ICRA).

[19] Tianshu Chu,et al. Multi-Agent Deep Reinforcement Learning for Large-Scale Traffic Signal Control , 2019, IEEE Transactions on Intelligent Transportation Systems.

[20] Saeid Nahavandi,et al. Deep Reinforcement Learning for Multiagent Systems: A Review of Challenges, Solutions, and Applications , 2018, IEEE Transactions on Cybernetics.

[21] Peng Ning,et al. Improving learning and adaptation in security games by exploiting information asymmetry , 2015, 2015 IEEE Conference on Computer Communications (INFOCOM).

[22] R Bellman,et al. On the Theory of Dynamic Programming. , 1952, Proceedings of the National Academy of Sciences of the United States of America.

[23] Jimmy Ba,et al. Adam: A Method for Stochastic Optimization , 2014, ICLR.

[24] Shimon Whiteson,et al. Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning , 2020, J. Mach. Learn. Res..

[25] David Silver,et al. Fictitious Self-Play in Extensive-Form Games , 2015, ICML.

[26] Jakub W. Pachocki,et al. Dota 2 with Large Scale Deep Reinforcement Learning , 2019, ArXiv.

[27] Hiroaki Kitano,et al. RoboCup: The Robot World Cup Initiative , 1997, AGENTS '97.

[28] Shane Legg,et al. Human-level control through deep reinforcement learning , 2015, Nature.

[29] Shimon Whiteson,et al. Counterfactual Multi-Agent Policy Gradients , 2017, AAAI.

[30] Jong-Hwan Kim,et al. AI World Cup: Robot-Soccer-Based Competitions , 2021, IEEE Transactions on Games.

[31] Etienne Perot,et al. Deep Reinforcement Learning framework for Autonomous Driving , 2017, Autonomous Vehicles and Machines.

[32] Marcin Andrychowicz,et al. Hindsight Experience Replay , 2017, NIPS.

[33] Taeyoung Kim,et al. Machine Learning for Advanced Wireless Sensor Networks: A Review , 2021, IEEE Sensors Journal.

[34] Yuval Tassa,et al. Continuous control with deep reinforcement learning , 2015, ICLR.

[35] Amnon Shashua,et al. Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving , 2016, ArXiv.

[36] Maxim Egorov Stanford. MULTI-AGENT DEEP REINFORCEMENT LEARNING , 2016 .

[37] Joonho Lee,et al. Learning agile and dynamic motor skills for legged robots , 2019, Science Robotics.

[38] Frans A. Oliehoek,et al. A Concise Introduction to Decentralized POMDPs , 2016, SpringerBriefs in Intelligent Systems.

[39] Peter Stone,et al. Deep Recurrent Q-Learning for Partially Observable MDPs , 2015, AAAI Fall Symposia.

[40] David Silver,et al. A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning , 2017, NIPS.

[41] Guy Lever,et al. Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward , 2018, AAMAS.

[42] Jürgen Schmidhuber,et al. Long Short-Term Memory , 1997, Neural Computation.

[43] Alex Graves,et al. Asynchronous Methods for Deep Reinforcement Learning , 2016, ICML.