Home /Research /Learning Scalable Deep Kernels with Recurrent\nStructure
OTHER

Learning Scalable Deep Kernels with Recurrent\nStructure

Maruan Al-Shedivat, Andrew Gordon Wilson, Yunus Saatchi, Zhiting Hu, Eric P. Xing

Year
2017
Citations
42

Abstract

Many applications in speech, robotics, finance, and biology deal with\nsequential data, where ordering matters and recurrent structures are common.\nHowever, this structure cannot be easily captured by standard kernel functions.\nTo model such structure, we propose expressive closed-form kernel functions for\nGaussian processes. The resulting model, GP-LSTM, fully encapsulates the\ninductive biases of long short-term memory (LSTM) recurrent networks, while\nretaining the non-parametric probabilistic advantages of Gaussian processes. We\nlearn the properties of the proposed kernels by optimizing the Gaussian process\nmarginal likelihood using a new provably convergent semi-stochastic gradient\nprocedure, and exploit the structure of these kernels for scalable training and\nprediction. This approach provides a practical representation for Bayesian\nLSTMs. We demonstrate state-of-the-art performance on several benchmarks, and\nthoroughly investigate a consequential autonomous driving application, where the\npredictive uncertainties provided by GP-LSTM are uniquely valuable.

Keywords

Computer scienceArtificial intelligenceKernel (algebra)Gaussian processRecurrent neural networkScalabilityMachine learningProbabilistic logicParametric statisticsRepresentation (politics)

Related papers

Browse all OTHER papers