Relaxed Attention for Transformer Models
Open paper
Benchmark
Model
Rank
Results
lipreading-on-lrs3-ted
AV-HuBERT Large + Relaxed Attention + LM
#7
Word Error Rate (WER): 25.51
machine-translation-on-iwslt2014-german
Cutoff + Relaxed Attention + LM
#3
BLEU score: 37.96
Number of Params: 24.1M
Rank counts only results with a code link.