Online TD(A) for discrete-time Markov jump linear systems
Rafael L. Beirigo, Marcos G. Todorov, André M. S. Barreto
- 发表年份
- 2018
- 引用次数
- 8
摘要
This paper proposes a new approach for the optimal quadratic control of discrete-time Markov jump linear systems (MJLS), inspired on the temporal differences (TD) concepts of reinforcement learning. The method is online, in the sense that it is able to simultaneously apply and refine the currently available controller, and it is transition model-free, because there is no need for explicit knowledge of the Markov chain transition probabilities, provided it can be sampled or simulated. The strategy builds upon a previously proposed offline method and we hope will pave the way for developing and adapting reinforcement learning techniques for MJLS. The method is experimentally evaluated in Samuelson's macroeconomic model and in the control of a faulty robotic manipulator arm, performing favorably when compared to its offline predecessor.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002