SLUE: New Benchmark Tasks for Spoken Language Understanding Evaluation on Natural Speech

Benchmark Model Rank Results
named-entity-recognition-on-slueW2V2-L-LL60K (pipeline approach, uses LM)#1F1 (%): 69.6label-F1 (%): 82.2Text model: DeBERTa-L
named-entity-recognition-on-slueW2V2-B-LS960 (pipeline approach, uses LM)#2F1 (%): 68.0label-F1 (%): 79.8Text model: DeBERTa-L
named-entity-recognition-on-slueW2V2-L-LL60K (e2e approach, uses LM)#4F1 (%): 64.8label-F1 (%): 73.3Text model: N/A
named-entity-recognition-on-slueW2V2-B-LS960 (e2e approach, uses LM)#5F1 (%): 63.4label-F1 (%): 71.7Text model: N/A
named-entity-recognition-on-slueHuBERT-B-LS960 (e2e approach, uses LM)#6F1 (%): 61.9label-F1 (%): 70.3Text model: N/A
named-entity-recognition-on-slueW2V2-B-VP100K (e2e approach, uses LM)#7F1 (%): 61.8label-F1 (%): 69.8Text model: N/A
named-entity-recognition-on-slueW2V2-L-LL60K (pipeline approach)#8F1 (%): 57.8label-F1 (%): 78.8Text model: DeBERTa-L
named-entity-recognition-on-slueW2V2-L-LL60K (e2e approach)#9F1 (%): 50.9label-F1 (%): 64.7
named-entity-recognition-on-slueW2V2-B-LS960 (e2e approach)#10F1 (%): 50.2label-F1 (%): 64.0
named-entity-recognition-on-slueHuBERT-B-LS960 (e2e approach)#11F1 (%): 49.8label-F1 (%): 62.9
named-entity-recognition-on-slueW2V2-B-LS960 (pipeline approach)#12F1 (%): 49.5label-F1 (%): 74.2Text model: DeBERTa-L
named-entity-recognition-on-slueW2V2-B-VP100K (e2e approach)#13F1 (%): 47.9label-F1 (%): 60.8
sentiment-analysis-on-slueW2V2-L-LL60K (pipeline approach, uses LM)#1Recall (%): 60.4F1 (%): 63.3Text model: DeBERTa-L
sentiment-analysis-on-slueW2V2-L-LL60K (pipeline approach)#2Recall (%): 60.2F1 (%): 63.3Text model: DeBERTa-L
sentiment-analysis-on-slueW2V2-B-LS960 (pipeline approach, uses LM)#3Recall (%): 60.0F1 (%): 62.9Text model: DeBERTa-L
sentiment-analysis-on-slueW2V2-B-LS960 (pipeline approach)#4Recall (%): 59.0F1 (%): 61.8Text model: DeBERTa-L
sentiment-analysis-on-slueW2V2-L-LL60K (e2e approach)#5Recall (%): 49.2F1 (%): 48.5Text model: N/A
sentiment-analysis-on-slueHuBERT-B-LS960 (e2e approach)#6Recall (%): 47.5F1 (%): 48.0Text model: N/A
sentiment-analysis-on-slueW2V2-B-LS960 (e2e approach)#7Recall (%): 46.0F1 (%): 46.6Text model: N/A
sentiment-analysis-on-slueW2V2-B-VP100K (e2e approach)#8Recall (%): 38.7F1 (%): 38.4Text model: N/A
speech-recognition-on-slueW2V2-L-LL60K (+ TED-LIUM 3 LM)#1VoxPopuli (Dev): 9.1VoxPopuli (Test): 9.3VoxCeleb (Dev): 9.1
speech-recognition-on-slueW2V2-B-LS960 (+ TED-LIUM 3 LM)#2VoxPopuli (Dev): 12.0VoxPopuli (Test): 12.2VoxCeleb (Dev): 13.2
speech-recognition-on-slueW2V2-L-LL60K (+ in-domain LM)#3VoxPopuli (Dev): 12.0VoxPopuli (Test): 12.5VoxCeleb (Dev): 11.8
speech-recognition-on-slueW2V2-L-LL60K#4VoxPopuli (Dev): 14.0VoxPopuli (Test): 12.1VoxCeleb (Dev): 11.0
speech-recognition-on-slueW2V2-B-LS960 (+ in-domain LM)#5VoxPopuli (Dev): 14.6VoxPopuli (Test): 15.2VoxCeleb (Dev): 15.2
speech-recognition-on-slueW2V2-B-LS960#6VoxPopuli (Dev): 17.2VoxPopuli (Test): 17.9VoxCeleb (Dev): 17.2
speech-recognition-on-slueHuBERT-B-LS960#7VoxPopuli (Dev): 18.6VoxPopuli (Test): 19.1VoxCeleb (Dev): 19.6
speech-recognition-on-slueW2V2-B-VP100K#8VoxPopuli (Dev): 21.6VoxPopuli (Test): 22.4VoxCeleb (Dev): 29.9