Effects of Hyper-Parameters for Deep Reinforcement Learning in Robotic Motion Mimicry: A Preliminary Study
Taewoo Kim, Joo-Haeng Lee
- 发表年份
- 2019
- 引用次数
- 4
摘要
When applying deep reinforcement learning to the motion mimicry problem between teacher and student robots, this paper reports the initial results of how various hyper-parameter configurations affect performance of learning processes and quality of generated motions. The hyperparameters considered in this study include the structure of policies such as convolutional and fully connected networks, the type of activation functions such as ReLU and hyperbolic tangent, and the number of input sequences such as one, four and eight. Under these deep neural network configurations, PPO reinforcement learning algorithm has been applied for learning. In the simulator environment, the teacher NAO robot demonstrates a target action repeatedly, and the learner NAO robot tries to learn that action. The target actions include handshaking and two-arm raising. Our experimental results show that fully connected networks outperform the convolutional counterparts both in training statistics and motion quality. For activation functions, however, we found an interesting mismatch between training and evaluation quality: for example, a configuration with higher rewards does not guarantee less motion discrepancy, which may suggest a new research direction to design better loss and reward functions for robotic motion mimicry.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002