A0C: Alpha Zero in Continuous Action Space

Thomas M. Moerland, Joost Broekens, Aske Plaat, Catholijn M. Jonker

发表年份: 2018
引用次数: 20
访问权限: 开放获取

摘要

A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go. These games have a discrete action space. However, many real-world reinforcement learning domains have continuous action spaces, for example in robotic control, navigation and self-driving cars. This paper presents the necessary theoretical extensions of Alpha Zero to deal with continuous action space. We also provide some preliminary experiments on the Pendulum swing-up task, empirically showing the feasibility of our approach. Thereby, this work provides a first step towards the application of iterated search and learning in domains with a continuous action space.

关键词

Action (physics)Reinforcement learningComputer scienceSpace (punctuation)NoveltyAlpha (finance)Zero (linguistics)InterleavingTask (project management)Artificial intelligence

A0C: Alpha Zero in Continuous Action Space

摘要

关键词

相关论文

Statistical Learning Theory

Artificial intelligence: a modern approach

Applied Nonlinear Control

A new optimizer using particle swarm theory