Path Integral Policy Improvement with Differential Dynamic Programming

Tom Lefebvre, Guillaume Crevecoeur

发表年份: 2019
引用次数: 7

摘要

Path Integral Policy Improvement with Covariance Matrix Adaptation (PI 2 -CMA) is a step-based model-free reinforcement learning approach that combines statistical estimation techniques with fundamental results from Stochastic Optimal Control. Basically, a policy distribution is improved iteratively using reward weighted averaging of the corresponding rollouts. It was assumed that PI 2 -CMA somehow exploited gradient information that was contained by the reward weighted statistics. To our knowledge we are the first to expose the principle of this gradient extraction rigorously. Our findings reveal that PI 2 -CMA essentially obtains gradient information similar to the forward and backward passes in the Differential Dynamic Programming (DDP) method. It is then straightforward to extend the analogy with DDP by introducing a feedback term in the policy update. This suggests a novel algorithm which we coin Path Integral Policy Improvement with Differential Dynamic Programming (PI 2 -DDP). The resulting algorithm is similar to the previously proposed Sampled Differential Dynamic Programming (SaDDP) but we derive the method independently as a generalization of the framework of PI 2 -CMA. Our derivations suggest to implement some small variations to SaDDP so to increase performance. We validated our claims on a robot trajectory learning task.

关键词

Dynamic programmingComputer sciencePath (computing)GeneralizationAlgorithmArtificial intelligenceMathematicsProgramming language

Path Integral Policy Improvement with Differential Dynamic Programming

摘要

关键词

相关论文

Statistical Learning Theory

Artificial intelligence: a modern approach

Fractional Differential Equations

Applied Nonlinear Control