Mask Attention Networks: Rethinking and Strengthen Transformer

Benchmark Model Rank Results
abstractive-text-summarization-on-cnn-dailyMask Attention Network#31ROUGE-1: 40.98ROUGE-2: 18.29ROUGE-L: 37.88
machine-translation-on-iwslt2014-germanMask Attention Network (small)#14BLEU score: 36.3Number of Params: 37M
machine-translation-on-wmt2014-english-germanMask Attention Network (big)#11BLEU score: 30.4Number of Params: 215M
machine-translation-on-wmt2014-english-germanMask Attention Network (base)#29BLEU score: 29.1Number of Params: 63M
text-summarization-on-gigawordMask Attention Network#18ROUGE-1: 38.28ROUGE-2: 19.46ROUGE-L: 35.46