Value iteration for continuous-state POMDPs
Josep M. Porta, Matthijs T. J. Spaan, Nikos Vlassis
- Year
- 2004
- Citations
- 10
- Access
- Open access
Abstract
We present a value iteration algorithm for learning to act in Partially Observable Markov Decision Processes (POMDPs) with continuous state spaces. Mainstream POMDP research focuses on the discrete case and this complicates its application to, e.g., robotic problems that are naturally modeled using continuous state spaces. The main difficulty in defining a (belief-based) POMDP in a continuous state space is that expected values over states must be defined using integrals that, in general, cannot be computed in closed from. In this report, we provide three main contributions to the literature on continuous-state POMDPs. First, we show that the optimal infinite-horizon value function over the continuous infinite dimensional POMDP belief space is piecewise linear and convex, and is defined by a finite set of supporting alpha-functions that are analogous to the alpha-vectors (hyperplanes) defining the value function of a discrete-state POMDP. Second, we show that, for a fairly general class of POMDP models in which all functions of interest are modeled by Gaussian mixtures, all belief updates and value iteration backups can be carried out analytically and exact. Contrary to the discrete case, in a continuous-state POMDP the alpha-functions may grow in size (e.g., in the number of Gaussian components) in each value iteration. Third, we show how the recent point-based value iteration algorithms for discrete POMDPs can be extended to the continuous case, allowing for efficient planning in practical problems. In particular, we demonstrate Perseus, our previously proposed randomized point-based value iteration algorithm, in a simple robot planning problem in a continuous domain, where encouraging results are observed.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991