Planning for human-robot interaction: representing time and human intention
Illah Nourbakhsh, Reid Simmons, Frank Broz
- Year
- 2008
- Citations
- 13
Abstract
This thesis proposes a novel approach to planning for a specific class of human-robot interaction domains: those in which robots engage in tasks with humans that are governed by social conventions. When humans perform these tasks, they try to achieve individual goals in an environment that they share with other people. Social conventions exist as a guideline for how to interact with others so that all parties involved can achieve their goals efficiently without interfering with one another. Recognizing what goals others are trying to achieve and performing actions at the appropriate time in the interaction are critical abilities for social competence. The approach to human-robot social interaction taken in this thesis focuses on creating more accurate models of social tasks for planning. Because the human participants are modeled as a part of the environment, the world state in these problems is dynamic and partially observable. Human intention is represented as hidden state in a partially observable Markov decision process (POMDP), and the time-dependence of action outcomes are explicitly modeled. A model structure designed by a human expert is combined with human task performance data. The resulting models are large and complex. State aggregation over the time dimension of the state space is used to trade off between the accuracy of the representation and its size in order to find sufficiently expressive models that can also be solved tractably. The utility of this approach is demonstrated by implementing a controller for a mobile robot that rides elevators with people and an agent in a driving simulator that performs the Pittsburgh left with human drivers. Performance is evaluated by comparing the policies obtained using the proposed modeling technique to policies developed using less expressive representations. In an interactions with human participants, the policies for time-dependent POMDP models with human intention as hidden state outperform the other policies, achieving both higher rewards and more positive evaluations for naturalness and social propriety of behavior.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002