Object Recognition from Short Videos for Robotic Perception

Ivan Bogun, Anelia Angelova, Navdeep Jaitly

发表年份: 2015
引用次数: 6
访问权限: 开放获取

摘要

Deep neural networks have become the primary learning technique for object recognition. Videos, unlike still images, are temporally coherent which makes the application of deep networks non-trivial. Here, we investigate how motion can aid object recognition in short videos. Our approach is based on Long Short-Term Memory (LSTM) deep networks. Unlike previous applications of LSTMs, we implement each gate as a convolution. We show that convolutional-based LSTM models are capable of learning motion dependencies and are able to improve the recognition accuracy when more frames in a sequence are available. We evaluate our approach on the Washington RGBD Object dataset and on the Washington RGBD Scenes dataset. Our approach outperforms deep nets applied to still images and sets a new state-of-the-art in this domain.

关键词

Artificial intelligenceComputer scienceDeep learningConvolutional neural networkObject (grammar)Convolution (computer science)Motion (physics)Computer visionPattern recognition (psychology)Domain (mathematical analysis)

Object Recognition from Short Videos for Robotic Perception

摘要

关键词

相关论文

Statistical Learning Theory

Artificial intelligence: a modern approach

Applied Nonlinear Control

A new optimizer using particle swarm theory