论文信息 - Fast Neural Machine Translation Implementation

Fast Neural Machine Translation Implementation

This paper describes the submissions to the efficiency track for GPUs at the Workshop for Neural Machine Translation and Generation by members of the University of Edinburgh, Adam Mickiewicz University, Tilde and University of Alicante. We focus on efficient implementation of the recurrent deep-learning model as implemented in Amun, the fast inference engine for neural machine translation. We improve the performance with an efficient mini-batching algorithm, and by fusing the softmax operation with the k-best extraction algorithm. Submissions using Amun were first, second and third fastest in the GPU efficiency track.

[1] Marcin Junczys-Dowmunt,et al. Is Neural Machine Translation Ready for Deployment? A Case Study on 30 Translation Directions , 2016, IWSLT.

[2] Steve Renals,et al. Multiplicative LSTM for sequence modelling , 2016, ICLR.

[3] Yoshua Bengio,et al. Neural Machine Translation by Jointly Learning to Align and Translate , 2014, ICLR.

[4] Rico Sennrich,et al. Nematus: a Toolkit for Neural Machine Translation , 2017, EACL.

[5] Kevin Skadron,et al. Enabling Task Parallelism in the CUDA Scheduler , 2009 .

[6] Rico Sennrich,et al. Neural Machine Translation of Rare Words with Subword Units , 2015, ACL.

[7] André F. T. Martins,et al. Marian: Fast Neural Machine Translation in C++ , 2018, ACL.

[8] Graham Neubig,et al. Findings of the Second Workshop on Neural Machine Translation and Generation , 2018, NMT@ACL.