MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Benchmark Model Rank Results
intent-classification-on-massivemT5 Base (encoder-only)#1Intent Accuracy: 86.1
intent-classification-on-massivemT5 Base (text-to-text)#2Intent Accuracy: 85.3
intent-classification-on-massiveXLM-R Base#3Intent Accuracy: 85.1
slot-filling-on-massiveXLM-R Base#1Slot F1 Score: 83.6
slot-filling-on-massivemT5 Base (encoder-only)#2Slot F1 Score: 82.2
slot-filling-on-massivemT5 Base (text-to-text)#3Slot F1 Score: 81.3
zero-shot-slot-filling-on-massiveXLM-R Base#1Slot F1 Score: 64.2
zero-shot-slot-filling-on-massivemT5 Base (encoder-only)#2Slot F1 Score: 56.9
zero-shot-slot-filling-on-massivemT5 Base (text-to-text)#3Slot F1 Score: 50.6