W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Benchmark Model Rank Results
speech-recognition-on-librispeech-test-cleanw2v-BERT XXL#29Word Error Rate (WER): 1.4
speech-recognition-on-librispeech-test-otherw2v-BERT XXL#10Word Error Rate (WER): 2.5