Cross-Class Relevance Learning for Temporal Concept Localization

2019-11-19 20:31:04

Junwei Ma, Satya Krishna Gorti, Maksims Volkovs, Ilya Stanevich, Guangwei Yu

arXiv_CV

arXiv_CV Video_Caption Classification Relation Pose Action

Abstract
Abstract (translated)
URL
PDF

Abstract

We present a novel Cross-Class Relevance Learning approach for the task of temporal concept localization. Most localization architectures rely on feature extraction layers followed by a classification layer which outputs class probabilities for each segment. However, in many real-world applications classes can exhibit complex relationships that are difficult to model with this architecture. In contrast, we propose to incorporate target class and class-related features as input, and learn a pairwise binary model to predict general segment to class relevance. This facilitates learning of shared information between classes, and allows for arbitrary class-specific feature engineering. We apply this approach to the 3rd YouTube-8M Video Understanding Challenge together with other leading models, and achieve first place out of over 280 teams. In this paper we describe our approach and show some empirical results.

Abstract (translated)

URL

https://arxiv.org/abs/1911.08548

PDF

https://arxiv.org/pdf/1911.08548.pdf