Home /Research /New concepts in team theory: mean field teams and reinforcement learning
SWARM

New concepts in team theory: mean field teams and reinforcement learning

Jalal Arabneydi

Year
2017
Citations
23
Access
Open access

Abstract

This thesis consists of two parts wherein each part introduces a new concept in team theory. In the first part, we introduce systems with partially exchangeable agents. A system is called partially exchangeable if it can be partitioned into sub-populations where agents are exchangeable. A sub-population of agents is called exchangeable if the manner in which agents are indexed does not affect the dynamics and cost. In practice, this insensitivity to the index naturally emerges in many applications. For example, in power systems, the system dynamics and cost would not change if the houses in a residential neighbourhood were numbered differently; in swarm robotics, the dynamics andcost depend on the position of the robots, not on how the agents are indexed. We first show that a system with partially exchangeable agents is equivalent to a system where agents are coupled in the dynamics and cost through the aggregate behaviour of agents (called mean-field). Then, we investigate and identify the optimal strategy---under mean-field sharing information structure which is non-classical---for two different models: linear quadratic and controlled Markov chain. We show that the optimal strategies, unlike the existing results in team theory, are scalable to large scale systems. We use the theory to solve idealized models of demand response in power systems and resource allocation in networks.In the second part, we study systems with partial history sharing information structure---which encompasses a large class of team problems including mean-field teams---when agents do not know the complete model of the system. The agents must learn the optimal strategies by interacting with their environment using reinforcement learning. We develop a reinforcement learning algorithm that guarantees epsilon-team-optimal performance. As an intermediate step of this development, we revisit the well-known partially observable Markov decision process and propose a novel approach to find an epsilon-optimal solution. The novelty of this approach is to identify the planning space based on the structure of the model. To illustrate the algorithm, we develop a reinforcement learning algorithm for the benchmark example of two-user multi access broadcast channel and present numerical results.

Keywords

Field (mathematics)Reinforcement learningLearning theoryArtificial intelligenceComputer sciencePsychologyMathematicsMathematics education

Related papers

Browse all SWARM papers