ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations

Benchmark Model Rank Results
lipreading-on-lrs2ES³ BaseWord Error Rate (WER): 30.7
lipreading-on-lrs2ES³ Base + extLMWord Error Rate (WER): 28.7
lipreading-on-lrs2ES³ Base*Word Error Rate (WER): 31.4
lipreading-on-lrs2ES³ Base* + extLMWord Error Rate (WER): 29.3
lipreading-on-lrs2ES³ LargeWord Error Rate (WER): 26.7
lipreading-on-lrs2ES³ Large + extLMWord Error Rate (WER): 24.6
lipreading-on-lrs3-tedES³ BaseWord Error Rate (WER): 40.3
lipreading-on-lrs3-tedES³ LargeWord Error Rate (WER): 37.1