Improving Neural Language Models by Segmenting, Attending, and Predicting the Future
Open paper
Benchmark
Model
Rank
Results
language-modelling-on-wikitext-103
Transformer-XL Large + Phrase Induction
#24
Test perplexity: 17.4
Number of params: 257M
Rank counts only results with a code link.