Generating Long Sequences with Sparse Transformers

Benchmark Model Rank Results
image-generation-on-imagenet-64x64Sparse Transformer 59M (strided)#34Bits per dim: 3.44
language-modelling-on-enwiki8Sparse Transformer (30 layers, fixed attn)#12Bit per Character (BPC): 0.99Number of params: 95M
open-domain-question-answering-on-searchqaSparse Attention#2EM: 64.7
question-answering-on-natural-questions-longSparse Attention#3F1: 74.5
question-answering-on-quasart-tSparse Attention#2EM: 52.1