Jointly Learning Visual and Auditory Speech Representations from Raw Data

Benchmark Model Rank Results
audio-visual-speech-recognition-on-lrs3-tedRAVEn Large#7Word Error Rate (WER): 1.4
lipreading-on-lrs2RAVEn Large#4Word Error Rate (WER): 18.6
lipreading-on-lrs3-tedRAVEn Large#5Word Error Rate (WER): 23.4
speech-recognition-on-lrs3-tedRAVEn Large#4Word Error Rate (WER): 1.4