Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Benchmark Model Rank Results
lipreading-on-lrs3-tedAV-HuBERT Large#9Word Error Rate (WER): 26.9
speech-recognition-on-lrs3-tedAV-HuBERT Large#3Word Error Rate (WER): 1.3