Amortized Noisy Channel Neural Machine Translation

2021-12-16 07:10:02

Richard Yuanzhe Pang, He He, Kyunghyun Cho

arXiv_CL

Abstract
Abstract (translated)
URL
PDF

Abstract

Noisy channel models have been especially effective in neural machine translation (NMT). However, recent approaches like "beam search and rerank" (BSR) incur significant computation overhead during inference, making real-world application infeasible. We aim to build an amortized noisy channel NMT model such that greedily decoding from it would generate translations that maximize the same reward as translations generated using BSR. We attempt three approaches: knowledge distillation, 1-step-deviation imitation learning, and Q learning. The first approach obtains the noisy channel signal from a pseudo-corpus, and the latter two approaches aim to optimize toward a noisy-channel MT reward directly. All three approaches speed up inference by 1-2 orders of magnitude. For all three approaches, the generated translations fail to achieve rewards comparable to BSR, but the translation quality approximated by BLEU is similar to the quality of BSR-produced translations.

Abstract (translated)

URL

https://arxiv.org/abs/2112.08670

PDF

https://arxiv.org/pdf/2112.08670.pdf