Evolving Multimodal Robot Behavior via Many Stepping Stones with the\n Combinatorial Multi-Objective Evolutionary Algorithm
Joost Huizinga, Jeff Clune
- 发表年份
- 2018
- 引用次数
- 3
- 访问权限
- 开放获取
摘要
An important challenge in reinforcement learning, including evolutionary\nrobotics, is to solve multimodal problems, where agents have to act in\nqualitatively different ways depending on the circumstances. Because multimodal\nproblems are often too difficult to solve directly, it is helpful to take\nadvantage of staging, where a difficult task is divided into simpler subtasks\nthat can serve as stepping stones for solving the overall problem.\nUnfortunately, choosing an effective ordering for these subtasks is difficult,\nand a poor ordering can reduce the speed and performance of the learning\nprocess. Here, we provide a thorough introduction and investigation of the\nCombinatorial Multi-Objective Evolutionary Algorithm (CMOEA), which avoids\nordering subtasks by allowing all combinations of subtasks to be explored\nsimultaneously. We compare CMOEA against two algorithms that can similarly\noptimize on multiple subtasks simultaneously: NSGA-II and Lexicase Selection.\nThe algorithms are tested on a multimodal robotics problem with six subtasks as\nwell as a maze navigation problem with a hundred subtasks. On these problems,\nCMOEA either outperforms or is competitive with the controls. Separately, we\nshow that adding a linear combination over all objectives can improve the\nability of NSGA-II to solve these multimodal problems. Lastly, we show that, in\ncontrast to NSGA-II and Lexicase Selection, CMOEA can effectively leverage\nsecondary objectives to achieve state-of-the-art results on the robotics task.\nIn general, our experiments suggest that CMOEA is a promising, state-of-the-art\nalgorithm for solving multimodal problems.\n
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002