Convolutional Tensor-Train LSTM for Spatio-temporal Learning

2020-02-21 05:00:01

Jiahao Su, Wonmin Byeon, Furong Huang, Jan Kautz, Animashree Anandkumar

arXiv_CV

arXiv_CV RNN CNN Relation Prediction Pose Action Video_Prediction

Abstract
Abstract (translated)
URL
PDF

Abstract

Higher-order Recurrent Neural Networks (RNNs) are effective for long-term forecasting since such architectures can model higher-order correlations and long-term dynamics more effectively. However, higher-order models are expensive and require exponentially more parameters and operations compared with their first-order counterparts. This problem is particularly pronounced in multidimensional data such as videos. To address this issue, we propose Convolutional Tensor-Train Decomposition (CTTD), a novel tensor decomposition with convolutional operations. With CTTD, we construct Convolutional Tensor-Train LSTM (Conv-TT-LSTM) to capture higher-order space-time correlations in videos. We demonstrate that the proposed model outperforms the conventional (first-order) Convolutional LSTM (ConvLSTM) as well as the state-of-the-art ConvLSTM-based approaches in pixel-level video prediction tasks on Moving-MNIST and KTH action datasets, but with much fewer parameters.

Abstract (translated)

URL

https://arxiv.org/abs/2002.09131

PDF

https://arxiv.org/pdf/2002.09131.pdf