Relaxed Attention for Transformer Models

Benchmark Model Rank Results
lipreading-on-lrs3-tedAV-HuBERT Large + Relaxed Attention + LM#7Word Error Rate (WER): 25.51
machine-translation-on-iwslt2014-germanCutoff + Relaxed Attention + LM#3BLEU score: 37.96Number of Params: 24.1M