Compressing Transformer-based self-supervised models for speech processing

2022-11-17 23:53:52

Tzu-Quan Lin, Tsung-Huan Yang, Chun-Yao Chang, Kuang-Ming Chen, Tzu-hsun Feng, Hung-yi Lee, Hao Tang

arXiv_SD

arXiv_SD Inference Knowledge Transformer Self-Supervised Speech

Abstract
Abstract (translated)
URL
PDF

Abstract

Despite the success of Transformers in self-supervised learning with applications to various downstream tasks, the computational cost of training and inference remains a major challenge for applying these models to a wide spectrum of devices. Several isolated attempts have been made to compress Transformers, prior to applying them to downstream tasks. In this work, we aim to provide context for the isolated results, studying several commonly used compression techniques, including weight pruning, head pruning, low-rank approximation, and knowledge distillation. We report wall-clock time, the number of parameters, and the number of multiply-accumulate operations for these techniques, charting the landscape of compressing Transformer-based self-supervised models.

Abstract (translated)

URL

https://arxiv.org/abs/2211.09949

PDF

https://arxiv.org/pdf/2211.09949.pdf

Compressing Transformer-based self-supervised models for speech processing

Abstract

Abstract (translated)

URL

PDF Copy

PDF