Hyperparameter Analysis for Image Captioning

Abstract
Abstract (translated)
URL
PDF

Abstract

In this paper, we perform a thorough sensitivity analysis on state-of-the-art image captioning approaches using two different architectures: CNN+LSTM and CNN+Transformer. Experiments were carried out using the Flickr8k dataset. The biggest takeaway from the experiments is that fine-tuning the CNN encoder outperforms the baseline and all other experiments carried out for both architectures.

Abstract (translated)

URL

https://arxiv.org/abs/2006.10923

PDF

https://arxiv.org/pdf/2006.10923.pdf