论文信息 - Exploiting structure in coordinating multiple decision makers

Exploiting structure in coordinating multiple decision makers

This thesis is concerned with sequential decision making by multiple agents, whether they are acting cooperatively to maximize team reward or selfishly trying to maximize their individual rewards. The practical intractability of this general problem led to efforts in identifying special cases that admit efficient computation, yet still represent a wide enough range of problems. In our work, we identify the class of problems with structured interactions, where actions of one agent can have non-local effects on the transitions and/or rewards of another agent. We addressed the following research questions: (1) How can we compactly represent this class of problems? (2) How can we efficiently calculate agent policies that maximize team reward (for cooperative agents) or achieve equilibrium (self-interested agents)? (3) How can we exploit structured interactions to make reasoning about communication offline tractable? For representing our class of problems, we developed a new decision-theoretic model, Event-Driven Interactions with Complex Rewards (EDI-CR), that explicitly represents structured interactions. EDI-CR is a compact yet general representation capable of capturing problems where the degree of coupling among agents ranges from complete independence to complete dependence. For calculating agent policies, we draw on several techniques from the field of mathematical optimization and adapt them to exploit the special structure in EDI-CR. We developed a Mixed Integer Linear Program formulation of EDI-CR with cooperative agents that results in programs much more compact and faster to solve than formulations ignoring structure. We also investigated the use of homotopy methods as an optimization technique, as well as formulation of self-interested EDI-CR as a system of non-linear equations. We looked at the issue of communication in both cooperative and self-interested settings. For the cooperative setting, we developed heuristics that assess the impact of potential communication points and add the ones with highest impact to the agents’ decision problems. Our heuristics successfully pick communication points that improve team reward while keeping problem size manageable. Also, by controlling the amount of communication introduced by a heuristic, our approach allows us to control the tradeoff between solution quality and problem size. For self-interested agents, we look at an example setting where communication is an integral part of problem solving, but where the self-interested agents have a reason to be reticent (e.g. privacy concerns). We formulate this problem as a game of incomplete information and present a general algorithm for calculating approximate equilibrium profile in this class of games.

Victor Lesser | Hala Mostafa | V. Lesser | Hala Mostafa

[1] Alain Dutech,et al. An Investigation into Mathematical Programming for Finite Horizon Decentralized POMDPs , 2014, J. Artif. Intell. Res..

[2] Tuomas Sandholm,et al. Finding equilibria in large sequential games of imperfect information , 2006, EC '06.

[3] D. Agrawal,et al. View Invalidation for Dynamic Content Caching in Multitiered Architectures , 2002, Very Large Data Bases Conference.

[4] Victor Lesser,et al. Exploiting structure in decentralized markov decision processes , 2006 .

[5] Victor R. Lesser,et al. Minimizing communication cost in a distributed Bayesian network using a decentralized MDP , 2003, AAMAS '03.

[6] P. Jean-Jacques Herings,et al. Computation of the Nash Equilibrium Selected by the Tracing Procedure in N-Person Games , 2002, Games Econ. Behav..

[7] Michael L. Littman,et al. Graphical Models for Game Theory , 2001, UAI.

[8] Kevin Leyton-Brown,et al. Computing Nash Equilibria of Action-Graph Games , 2004, UAI.

[9] Kevin Leyton-Brown,et al. Temporal Action-Graph Games: A New Representation for Dynamic Games , 2009, UAI.

[10] Victor R. Lesser,et al. Multi-agent policies: from centralized ones to decentralized ones , 2002, AAMAS '02.

[11] Bruce M. Maggs,et al. Invalidation Clues for Database Scalability Services , 2007, 2007 IEEE 23rd International Conference on Data Engineering.

[12] Victor R. Lesser,et al. Agent interaction in distributed POMDPs and its implications on complexity , 2006, AAMAS '06.

[13] Eitan Altman,et al. Zero-sum constrained stochastic games with independent state processes , 2005, Math. Methods Oper. Res..

[14] Shlomo Zilberstein,et al. Dynamic Programming for Partially Observable Stochastic Games , 2004, AAAI.

[15] Shimon Whiteson,et al. Lossless clustering of histories in decentralized POMDPs , 2009, AAMAS.

[16] Masha Sosonkina,et al. Algorithm 777: HOMPACK90: a suite of Fortran 90 codes for globally convergent homotopy algorithms , 1997, TOMS.

[17] Claudia V. Goldman,et al. Optimizing information exchange in cooperative multi-agent systems , 2003, AAMAS '03.

[18] Victor R. Lesser,et al. Self-interested database managers playing the view maintenance game , 2008, AAMAS.

[19] R. McKelvey,et al. Computation of equilibria in finite games , 1996 .

[20] Shlomo Zilberstein,et al. Formal models and algorithms for decentralized decision making under uncertainty , 2008, Autonomous Agents and Multi-Agent Systems.

[21] Manuela M. Veloso,et al. Reasoning about joint beliefs for execution-time communication decisions , 2005, AAMAS '05.

[22] Robert Wilson,et al. A global Newton method to compute Nash equilibria , 2003, J. Econ. Theory.

[23] Victor R. Lesser,et al. Compact Mathematical Programs For DEC-MDPs With Structured Agent Interactions , 2011, UAI.

[24] Daphne Koller,et al. Multi-Agent Influence Diagrams for Representing and Solving Games , 2001, IJCAI.

[25] Victor R. Lesser,et al. Decentralized Markov decision processes with event-driven interactions , 2004, Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems, 2004. AAMAS 2004..

[26] Daphne Koller,et al. Multi-agent algorithms for solving graphical games , 2002, AAAI/IAAI.

[27] Edmund H. Durfee,et al. Commitment-driven distributed joint policy search , 2007, AAMAS '07.

[28] Pierfrancesco La Mura. Game Networks , 2000, UAI.

[29] Daphne Koller,et al. A Continuation Method for Nash Equilibria in Structured Games , 2003, IJCAI.

[30] Vincent Conitzer,et al. Complexity Results about Nash Equilibria , 2002, IJCAI.

[31] Qiong Luo,et al. Template-Based Runtime Invalidation for Database-Generated Web Contents , 2004, APWeb.

[32] Ronald A. Howard,et al. Influence Diagrams , 2005, Decis. Anal..

[33] Layne T. Watson,et al. Probability-one homotopy maps for mixed complementarity problems , 2008, Comput. Optim. Appl..

[34] Francisco S. Melo,et al. Interaction-driven Markov games for decentralized multiagent planning under uncertainty , 2008, AAMAS.

[35] Makoto Yokoo,et al. Communications for improving policy computation in distributed POMDPs , 2004, Proceedings of the Third International Joint Conference on Autonomous Agents and Multiagent Systems, 2004. AAMAS 2004..

[36] H. Kuk. On equilibrium points in bimatrix games , 1996 .

[37] Miroslav Dudík,et al. A Sampling-Based Approach to Computing Equilibria in Succinct Extensive-Form Games , 2009, UAI.

[38] R. McKelvey,et al. Quantal Response Equilibria for Extensive Form Games , 1998 .

[39] Daniel L. Silver,et al. A distributed multi-agent meeting scheduler , 2008, J. Comput. Syst. Sci..

[40] B. Stengel,et al. Efficient Computation of Behavior Strategies , 1996 .

[41] Y. Saad,et al. GMRES: a generalized minimal residual algorithm for solving nonsymmetric linear systems , 1986 .

[42] Shlomo Zilberstein,et al. Memory-Bounded Dynamic Programming for DEC-POMDPs , 2007, IJCAI.

[43] Ronald A. Howard,et al. Readings on the Principles and Applications of Decision Analysis , 1989 .

[44] Shlomo Zilberstein,et al. Improved Memory-Bounded Dynamic Programming for Decentralized POMDPs , 2007, UAI.

[45] Nikos A. Vlassis,et al. Decentralized planning under uncertainty for teams of communicating agents , 2006, AAMAS '06.

[46] Leslie Pack Kaelbling,et al. Multi-Agent Filtering with Infinitely Nested Beliefs , 2008, NIPS.

[47] François Charpillet,et al. Point-based Dynamic Programming for DEC-POMDPs , 2006, AAAI.