首页 /研究 /Multiagent Reinforcement Learning with Regret Matching for Robot Soccer
LEARNING

Multiagent Reinforcement Learning with Regret Matching for Robot Soccer

Qiang Liu, Jiachen Ma, Wei Xie

发表年份
2013
引用次数
5
访问权限
开放获取

摘要

This paper proposes a novel multiagent reinforcement learning (MARL) algorithm Nash-<svg style="vertical-align:-1.90608pt;width:13.175px;" id="M1" height="14.3125" version="1.1" viewBox="0 0 13.175 14.3125" width="13.175" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns="http://www.w3.org/2000/svg"> <g transform="matrix(.017,-0,0,-.017,.062,11.4)"><path id="x1D444" d="M745 361q0 -134 -83.5 -233t-214.5 -130l16 -11q97 -67 250 -132l-8 -23q-76 3 -131 16q-81 19 -242 125l-20 13q-129 8 -209 91t-80 208q0 160 116 271t289 111q136 0 226.5 -83t90.5 -223zM645 356q0 127 -57.5 201.5t-169.5 74.5q-126 0 -210.5 -104.5t-84.5 -248.5&#xA;q0 -97 46 -166.5t129 -87.5l84 15l29 -19q104 21 169 121.5t65 213.5z" /></g> </svg> learning with regret matching, in which regret matching is used to speed up the well-known MARL algorithm Nash-<svg style="vertical-align:-1.90608pt;width:13.175px;" id="M2" height="14.3125" version="1.1" viewBox="0 0 13.175 14.3125" width="13.175" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns="http://www.w3.org/2000/svg"> <g transform="matrix(.017,-0,0,-.017,.062,11.4)"><use xlink:href="#x1D444"/></g> </svg> learning. It is critical that choosing a suitable strategy for action selection to harmonize the relation between exploration and exploitation to enhance the ability of online learning for Nash-<svg style="vertical-align:-1.90608pt;width:13.175px;" id="M3" height="14.3125" version="1.1" viewBox="0 0 13.175 14.3125" width="13.175" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns="http://www.w3.org/2000/svg"> <g transform="matrix(.017,-0,0,-.017,.062,11.4)"><use xlink:href="#x1D444"/></g> </svg> learning. In Markov Game the joint action of agents adopting regret matching algorithm can converge to a group of points of no-regret that can be viewed as coarse correlated equilibrium which includes Nash equilibrium in essence. It is can be inferred that regret matching can guide exploration of the state-action space so that the rate of convergence of Nash-<svg style="vertical-align:-1.90608pt;width:13.175px;" id="M4" height="14.3125" version="1.1" viewBox="0 0 13.175 14.3125" width="13.175" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns="http://www.w3.org/2000/svg"> <g transform="matrix(.017,-0,0,-.017,.062,11.4)"><use xlink:href="#x1D444"/></g> </svg> learning algorithm can be increased. Simulation results on robot soccer validate that compared to original Nash-<svg style="vertical-align:-1.90608pt;width:13.175px;" id="M5" height="14.3125" version="1.1" viewBox="0 0 13.175 14.3125" width="13.175" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns="http://www.w3.org/2000/svg"> <g transform="matrix(.017,-0,0,-.017,.062,11.4)"><use xlink:href="#x1D444"/></g> </svg> learning algorithm, the use of regret matching during the learning phase of Nash-<svg style="vertical-align:-1.90608pt;width:13.175px;" id="M6" height="14.3125" version="1.1" viewBox="0 0 13.175 14.3125" width="13.175" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns="http://www.w3.org/2000/svg"> <g transform="matrix(.017,-0,0,-.017,.062,11.4)"><use xlink:href="#x1D444"/></g> </svg> learning has excellent ability of online learning and results in significant performance in terms of scores, average reward and policy convergence.

关键词

RegretScalable Vector GraphicsReinforcement learningMatching (statistics)Artificial intelligenceComputer scienceMathematicsCombinatoricsMachine learningStatistics

相关论文

查看 LEARNING 分类全部论文