Deep Residual Output Layers for Neural Language Generation

Benchmark Model Rank Results
language-modelling-on-penn-treebank-wordAWD-LSTM-DRILL + dynamic eval#11Test perplexity: 49.4Validation perplexity: 49.5Params: 24M
language-modelling-on-penn-treebank-wordAWD-LSTM-DRILL#24Test perplexity: 55.7Validation perplexity: 58.2Params: 24M
language-modelling-on-wikitext-2AWD-LSTM-DRILL + dynamic eval#16Test perplexity: 42.0Validation perplexity: 43.9Number of params: 34M
language-modelling-on-wikitext-2AWD-LSTM-DRILL#27Test perplexity: 61.9Validation perplexity: 64.9Number of params: 34M
machine-translation-on-wmt2014-english-germanTransformer-DRILL Base#41BLEU score: 28.1