Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Open paper
Benchmark
Model
Rank
Results
lipreading-on-lrs3-ted
AV-HuBERT Large
#9
Word Error Rate (WER): 26.9
speech-recognition-on-lrs3-ted
AV-HuBERT Large
#3
Word Error Rate (WER): 1.3
Rank counts only results with a code link.