RoBERTuito: a pre-trained language model for social media text in Spanish

Abstract
Abstract (translated)
URL
PDF

Abstract

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training, specially-crafted models for particular domains, such as scientific papers, medical documents, and others. In this work, we present RoBERTuito, a pre-trained language model for user-generated content in Spanish. We trained RoBERTuito on 500 million tweets in Spanish. Experiments on a benchmark of 4 tasks involving user-generated text showed that RoBERTuito outperformed other pre-trained language models for Spanish. In order to help further research, we make RoBERTuito publicly available at the HuggingFace model hub.

Abstract (translated)

URL

https://arxiv.org/abs/2111.09453

PDF

https://arxiv.org/pdf/2111.09453.pdf