GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Benchmark Model Rank Results
common-sense-reasoning-on-arc-challengeGLaM 64B/64E (0 shot)Accuracy: 50.3
common-sense-reasoning-on-arc-challengeGLaM 64B/64E (1 shot)Accuracy: 48.2
common-sense-reasoning-on-arc-easyGLaM (64B/64E) (5-shot)Accuracy: 74.8
common-sense-reasoning-on-arc-easyGLaM 64B/64E (0-shot)Accuracy: 68.0
language-modelling-on-lambadaGLaM 62B/64E (One-Shot)Accuracy: 80.9
question-answering-on-natural-questionsGLaM 62B/64E (Few-Shot)EM: 32.5
question-answering-on-natural-questionsGLaM 62B/64E (One-Shot)EM: 26.3
question-answering-on-natural-questionsGLaM 62B/64E (Zero-Shot)EM: 24.7
question-answering-on-triviaqaGLaM 62B/64E (Few-shot)EM: 75.8
question-answering-on-triviaqaGLaM 62B/64E (One-shot)EM: 75.8
question-answering-on-triviaqaGLaM 62B/64E (Zero-shot)EM: 71.3
question-answering-on-webquestionsGLaM 62B/64E (Zero-Shot)EM: 15.5