| Benchmark | Model | Rank | Results |
|---|---|---|---|
| image-generation-on-imagenet-64x64 | Sparse Transformer 59M (strided) | #34 | Bits per dim: 3.44 |
| language-modelling-on-enwiki8 | Sparse Transformer (30 layers, fixed attn) | #12 | Bit per Character (BPC): 0.99Number of params: 95M |
| open-domain-question-answering-on-searchqa | Sparse Attention | #2 | EM: 64.7 |
| question-answering-on-natural-questions-long | Sparse Attention | #3 | F1: 74.5 |
| question-answering-on-quasart-t | Sparse Attention | #2 | EM: 52.1 |