Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
Open paper
Benchmark
Model
Rank
Results
keyword-spotting-on-google-speech-commands
M2D
#24
Google Speech Commands V2 35: 98.5
speaker-identification-on-voxceleb1
M2D ratio=0.6
#5
Top-1 (%): 94.8
Accuracy: 94.8
Rank counts only results with a code link.