Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input

Benchmark Model Rank Results
keyword-spotting-on-google-speech-commandsM2D#24Google Speech Commands V2 35: 98.5
speaker-identification-on-voxceleb1M2D ratio=0.6#5Top-1 (%): 94.8Accuracy: 94.8