Multimodal End-to-End Sparse Model for Emotion Recognition

2021-03-17 14:05:05

Wenliang Dai, Samuel Cahyawijaya, Zihan Liu, Pascale Fung

arXiv_CL

arXiv_CL Recognition Attention Sparse Action Emotion

Abstract
Abstract (translated)
URL
PDF

Abstract

Existing works on multimodal affective computing tasks, such as emotion recognition, generally adopt a two-phase pipeline, first extracting feature representations for each single modality with hand-crafted algorithms and then performing end-to-end learning with the extracted features. However, the extracted features are fixed and cannot be further fine-tuned on different target tasks, and manually finding feature extraction algorithms does not generalize or scale well to different tasks, which can lead to sub-optimal performance. In this paper, we develop a fully end-to-end model that connects the two phases and optimizes them jointly. In addition, we restructure the current datasets to enable the fully end-to-end training. Furthermore, to reduce the computational overhead brought by the end-to-end model, we introduce a sparse cross-modal attention mechanism for the feature extraction. Experimental results show that our fully end-to-end model significantly surpasses the current state-of-the-art models based on the two-phase pipeline. Moreover, by adding the sparse cross-modal attention, our model can maintain performance with around half the computation in the feature extraction part.

Abstract (translated)

URL

https://arxiv.org/abs/2103.09666

PDF

https://arxiv.org/pdf/2103.09666.pdf