Automated Audio Captioning with Epochal Difficult Captions for Curriculum Learning

2022-06-04 06:42:05

Andrew Koh, Soham Tiwari, Chng Eng Siong

arXiv_CL

Abstract
Abstract (translated)
URL
PDF

Abstract

In this paper, we propose an algorithm, Epochal Difficult Captions, to supplement the training of any model for the Automated Audio Captioning task. Epochal Difficult Captions is an elegant evolution to the keyword estimation task that previous work have used to train the encoder of the AAC model. Epochal Difficult Captions modifies the target captions based on a curriculum and a difficulty level determined as a function of current epoch. Epochal Difficult Captions can be used with any model architecture and is a lightweight function that does not increase training time. We test our results on three systems and show that using Epochal Difficult Captions consistently improves performance

Abstract (translated)

URL

https://arxiv.org/abs/2206.01918

PDF

https://arxiv.org/pdf/2206.01918.pdf