Visually Guided Self Supervised Learning of Speech Representations

Benchmark Model Rank Results
speech-emotion-recognition-on-crema-dGRUAccuracy: 55.01