A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Benchmark Model Rank Results
intent-classification-on-slurpPartially Fine-tuned HuBERTAccuracy (%): 87.51
slot-filling-on-slurpPartially Fine-tuned HuBERTF1: 0.753
speaker-verification-on-voxceleb1Fine-tuned HuBERT LargeEER: 2.36
speech-emotion-recognition-on-iemocapPartially Fine-tuned HuBERT LargeWA CV: 0.730WA: 0.796